Skip to content

Sep 12, 2024

Solana SVM and TPU

Notes on Solana SVM and TPU

SVM - What is it?

  • Register-based VM built for parallel execution
  • Runs programs compiled to Berkley Packet filter (BPF) code.
  • Each validator node in the network process the same txs in isolation - ensuring deterministic changes thus, reaching consensus.
  • Each transaction updates the global status, thus all validators must run them to keep their local state copy updated

Account-based data model

  • It uses a account-based data model -> All data lives in accounts.
  • All txs declare in advance which accounts they will read or write from -> Knowing which accounts can be accessed, enables parallelism of non-conflicting txs (txs that don't read/write from the same account) and simultaneous exeuction of conflicting txs to avoiud race conditions or double spend

Sealevel

  • Sealevel is SVM's parallel execution engine -> It schedules multiple ixs across differnt CPU cores as long as they are not conflicting -> Txs with no overlapping accounts run in parallel.
  • This multi-threaded VM approach of solana enales utilization of CPUs and GPUs even, getting to higher TPS and lower latency as it removes the bottleneck of sequential execution

Runtime

  • Each program is sotred in a read-only program account
  • Mutable state for apps (token balances, orderbooks, etc) lives in another account.
  • SVM enforces write locks on accounts to prevent conflicts
  • Solana's Cloudbreak is the storage engine that makes possible for validatgors to read/write multiple accounts in parallel
  • It is memory-mapped file based with concurrent datga structures.
  • On execution programs run on sandboxed BPF VM with a defined compute budget to prevent txs from hogging resources.-> If tx exeeds CU budget, it aborts.

TPU - What is it?

  • Transaction Processing Unit -> Software specialized pipeline for handling incoming transactions.
  • Block producing engine of the leader validator
  • Ingest, verify, execute and disseminate tx
  • Inspired by CPU pipeline design.
  • TPU is pipelined and paralelized such taht while some tx are being executed locally, the next ones are already being fetched and the old ones are already being packaged and broadcasted.
  • Stages of TPU: Fetch, SigVerify, Banking, PoH Recording, Shredding and Broadcasting

TPU-stages

Fetch Stage

  • Receive inbound txs from the network
  • Listen to UDP ports and sockets
  • Separte sockets for tpu - the main port for normal txs; tpu_vote - incoming vote txs from validators; tpu_forwards - port for forwarded txs that a previous leader could not handle
  • Separtion prevents high tx volumne from delaying critical votes.
  • This stage pull packages from the network interface and batches them in packages of 128 for efficiency and enqueues these packages for the next stage.
  • In this stage, txs are just raw packages, software has not yet parsed or verified them
  • The goal is to ingest data as fast as possible, keeping the network card ready and avoiding packet loss.

SigVerify

  • In charge of cryptographic signature checking.
  • Computationally expensive and main bottleneck -> Solana parallelizies it using GPU for bulk check
  • Validator just checks that the signatures are valid
  • Includes DOS protection, that would drop packets by IP to shed load
  • txs and votes are not processed inthe same lane so that voting packages do not get starved.
  • When this stage ends, the packages are guaranteeed to be from legitimate senders.
  • Use CPU and GPU to avoid bottlenecks.

Banking stages

  • Txs executed against the current leader state.
  • Once a batch of txs have legitimate singatures, they enter the banking (runtime execution engine) stage.
  • Multiple worker threads spawn to handle parallel tx processing -> 4 threads for non-voting txs and 2 for voting txs
  • Each thread pops tx from the queue and filters them by QoS (priotizing fee or age) and ensuring txs are not conflicting ones.
  • Up to 64 txs can be collected and reordered
  • There is serial execution per instruction, and even if one tx if invokes a program, the other threads can still run txs on a disjoint state.
  • If an ix within a ix fails, it is abnorted and effects are discarded.
  • If the tx success, the state changes are comitted to the in-memory ledger called Bank
  • In this stage is where the state is actually updated
  • This stage is CPU bound, meaning execution happens at CPU cores and it is optimized to maximize utilization of all cores by dividing the work across all threads.
  • By the time the badge is processed, all trasactions have either been aborted or been executed successfully (state updated)
  • At this point the txs are ready to be recorded in the ledger.