The main question of this lecture: a brand-new model's first working build decodes at 35 tokens/s at batch size 1 on four GB300s. Days later the same server does 873. Same GPU, same math, same weights. Which interventions made up those 25×, in what order were they found, and which ones do you already know how to apply to the next model you are asked to serve?
This lecture is a case study, and case studies are where the course's abstractions pay rent. Lectures 1–9 built the tools: the roofline and the 2P rule (L1), the KV cache as the center of the system (L2), speculative decoding (L8), and the kernel level: launches, graphs, fusion (L7). In September 2026 the SGLang team and friends got to use all of them at once, because day-0 support for DeepSeek-V4.1 Flash was not just "add a config": the model arrived with a new encoder–decoder backbone, a KV cache four times smaller than its predecessor, a hierarchical sparse-attention indexer, a four-stream residual connector, an n-gram memory, and, in the official checkpoint, its own speculative-decoding drafts. Everything below is measured throughput at BS=1 on 4× GB300 (batch size 1, one request at a time, on four NVIDIA GB300 GPUs), the setting in which every wasted microsecond is visible: no batch to hide latency behind, no prefill to amortize launches over. It is the purest exam a serving stack can take.
Before the ladder, the model itself. Every bar you will meet later exists because of something in this section, so it is worth going slowly.
V4.1 Flash is built on one bet: grow what makes the model smarter, shrink what every step pays for. And you already know what a step pays for. From Lecture 1, at batch size 1 a decode step reads the weights it will use and the stored notes (KV) from past tokens, and does almost no arithmetic on them. So two numbers decide how fast a model can decode: active parameters per token (how much weight a step must read) and KV bytes per token (how much state it must store and revisit). The table below is the whole model reduced to the rows a serving engineer should look at first:
| V4 Flash | V4.1 Flash | |
|---|---|---|
| Backbone | 284B | 552B, plus 196B of Engram memory parameters |
| Active per token | ~13B | ~8B on input, ~16B on output |
| Structure | 43 decoder layers | 20 causal encoder layers + 20 decoder layers |
| Global attention | CSA / HCA (compressor + hybrid attention) | CSA2: KV and indices shared across layers |
| Residual connector | single stream | mHC: four streams, Sinkhorn-mixed |
| Speculative drafts | none | DSpark: 3 draft blocks, block size 5 |
| Global KV per token | 3,514 bytes | 890 bytes |
Row by row: the totals went up (552B plus 196B of Engram memory, against 284B), which is where the smarts come from. What a step pays tells a different story: input tokens activate ~8B parameters instead of ~13B, output tokens ~16B, and the memory each token leaves behind fell fourfold, from 3,514 to 890 bytes. This is not a cheaper step; it is a reshaped one: cheaper reading, unchanged writing, a quarter of the memory, all wrapped in machinery no serving stack had kernels for yet. Hence the day-0 crawl. Four of those new pieces deserve a proper look now. (The other two, the mHC connector and the DSpark drafts, get their own sections: Skills 3 and 5.)
New piece one: a reader and a writer. Serving a long prompt involves two jobs: reading the prompt, then writing the answer one word at a time. V4.1 gives the two jobs to different layers (the third row of the table). The first 20 layers form the reader, a causal encoder, which processes the prompt with ordinary self-attention. "causal" is doing quiet but essential work here: it means each token's KV behaves exactly as in Lecture 2, so the usual cache machinery applies unchanged. And when the reading is done, the encoder does not hand down twenty layers of state. It compresses its final output into a single global KV cache.
The other 20 layers form the writer, the decoder. The honest way to see what changed is a before-and-now of one decode step, for one writer layer:
| one writer layer, one decode step | before (V4 Flash) | now (V4.1 Flash) |
|---|---|---|
| its question (Q) | built fresh from its own current state, used once, discarded | the same, unchanged: never shared, never cached, in any model |
| its notes about the prompt (global K/V) | written and stored at that layer | not written at all; the layer reads one shared bank |
| its notes about recent context (local K/V) | its own sliding-window KV, persisted across the prompt | its own local window, rebuilt on demand from the last 128 tokens |
The question row is the point, precisely because it is identical in both columns. A query is a one-step question built from a layer's current thought, and layer 22's thought already contains layer 21's answer, so sharing Q was never an option and caching it never useful. The real novelty is the middle row: who writes the notes. That settles the worry the shared bank raises on first hearing: twenty layers reading one notebook need not attend identically, because each still asks its own question. Same notebook, different questions, different answers. (Credit: YoCo, 2024, first let an upper half of a model share its lower half's KV.)
Put a concrete prompt through it: of the 40 layers a 4,096-token prompt passes through, only about half do deep careful attention, and the rest read the shared summary and the replayed 128-token tail. In complexity terms, writing \(N\) for the prompt length and \(L\) for the 40 layers, prefill drops from \(O(NL)\) to \(O(NL/2 + n_{\text{win}} \cdot L/2)\), where \(n_{\text{win}} = 128\) is the local window. Halved, with a small tail tax. "Roughly", because the re-runs are real work, not a free lunch.
Now the twist this lecture hangs on. Decode cannot take the shortcut. Every newly written token enters at the reader's first layer and leaves at the writer's last: all 40 layers, every step, on a model whose per-step weight read did not shrink. The architecture gave prefill a one-time discount per request, and gave decode nothing. That is why everything from here on is measured on decode.
New piece two: the shared notes, counted out like coins. The 890 bytes are the number Lecture 2 trained you to stare at, so let us count them. But first, one step back: what deserves storing at all? Watch a new token arrive. It comes with its own fresh query Q, and once that query has scored the past, its job is done forever: every future token will bring a new Q of its own. What all future tokens keep needing is the old positions' K and V. That is why the cache holds K and V only. The name is literal.
Put it in one line, with the symbols spelled out (Lecture 2 derived this slowly and we only borrow the result). At position \(t\) the model builds a query \(q_t\), a vector of length \(d\) (the attention width); it matches that query against the stored key vectors \(K_{\le t}\) of every position up to \(t\), scaled by \(\sqrt d\) so the scores stay tame; the softmax turns matches into weights; and the weights mix the stored value vectors \(V_{\le t}\) into the layer's output:
\[ \text{output}_t \;=\; \mathrm{softmax}\!\left( \frac{q_t\, K_{\le t}^{\top}}{\sqrt{d}} \right)\, V_{\le t} \]The \(q_t\) is discarded the moment the step ends; the \(K_{\le t}\) and \(V_{\le t}\) under it only ever grow. Which brings us back to the question that makes V4.1 interesting: what exactly is written into that growing memory, and by whom?
Now the rules, three of them. Rule 1: share the notes. CSA2 statically assigns every global-attention layer one of three modes. Full: compute main KV, run the indexer, emit a fresh Top-K selection. Reindex: write nothing; rescan the most recent Full layer's bank with your own indexer query, hand a fresh selection to your group. Reuse: write nothing, pick nothing; read the bank with the latest selection. The layout is strikingly regular: the encoder's 18 global-attention layers (its first two are pure sliding window) form three groups of six, each headed by a Full layer; the decoder's 20 form five groups of four, headed Full, Reindex, Reindex, Reindex, Reindex. So four layers write global KV at all: encoder layers 2, 8 and 14 (zero-indexed) plus the first decoder layer, whose bank is projected from the encoder's output with its own weights. The other 36 keep no global state. Note what sharing does not do: it does not make layers look alike. Each computes its own Q, and each Reindex head re-scores the shared bank to pick fresh rows for its group.
Rule 2: merge pairs of tokens. Of the four note-taking layers, the first three merge every two adjacent tokens into a single entry, so each of them costs only half an entry per original token: \(3\times\tfrac12 = \tfrac32\). The fourth keeps full resolution, one entry per token. Total: 2.5 entries per token.
Rule 3: write small. One entry is 288 bytes of main KV plus 68 bytes of indexer K. The 288 counts out easily: 512 latent channels at 4 bits (E2M1 codes: 1 sign bit, 2 exponent bits, 1 mantissa bit) are 256 bytes, plus 32 one-byte E4M3 scales (8-bit floats), one per 16 channels. The values allow this: after RMSNorm and RoPE (which rotates channels pairwise and so preserves norms), the latent's norm stays under about \(\sqrt{512} \approx 22.6\), and if the whole vector is under that bound, no single channel can pass it either; training never saw one above 10. The cache is quantized after RoPE, on purpose. The sliding-window KV, more fragile, stays in FP8. Last, the singular in "main KV" is real: in the reference attention kernel one selected vector both scores the query and feeds the output sum; K and V are one array, not two. Put the three rules together:
\[ \underbrace{(288 + 68)}_{\text{one entry: codes + scales}}\ \times\ \left(\! \underbrace{3 \times \tfrac{1}{2}}_{\text{3 writers, pooled}}\ +\ \underbrace{1}_{\text{1 writer, full res}}\!\right)\ =\ \underbrace{890\ \text{bytes}}_{\frac{1}{4}\text{ of V4 Flash}} \]Two receipts, since you can check both of them in the code. One entry charges half a byte per channel for the codes plus one scale byte per channel group. For the main KV: 512 channels, one scale per 16 channels. For the indexer's K: 128 channels, one scale per 32 channels:
\[ 288 \;=\; \underbrace{512 \times \tfrac12}_{\text{e2m1 codes}}\ +\ \underbrace{\tfrac{512}{16}}_{\text{e4m3 scales}} \qquad\qquad 68 \;=\; \underbrace{128 \times \tfrac12}_{\text{e2m1 codes}}\ +\ \underbrace{\tfrac{128}{32}}_{\text{e4m3 scales}} \](The bill counts only the global bank. SGLang repeats the receipts in code: the fp4 layout in
c1.cuh, and the indexer size computed as index_head_dim//2 + index_head_dim//32 =
64 + 4 in deepseek_v4_memory_pool.py.)
In words: four layers take notes, three of them take one note per two tokens, and everything is written in tiny 4-bit type. A quarter of V4 Flash's 3,514 bytes, and here it is to scale:
Three things follow. First, capacity: with four times smaller notes, the same memory room holds
four times the history. That is why the launch command late in this lecture can afford
--max-total-tokens 33554432; a 33-million-token pool only fits because one token
costs 890 bytes and not 3,514.
Capacity again, one level down. The 890 bytes live in HBM. The persistent cache that survives sessions lives on SSD, and there the saving compounds to an eighth: the same quarter-size global KV, and no sliding-window KV at all; SWA state lives for minutes in a host-DRAM pool, reused only inside an active session. An expired window entry costs exactly one replay of the last 128 tokens: the report's "SWA Bounded Replay", approximate but trained for.
Second, notice where the saving came from. Not from clever number formats: mostly from deciding that 36 layers should keep nothing. Sharing did the heavy lifting, merging trimmed a further third (\(4 \to 2.5\) entries), and FP4 halved the entry itself on top. Compression was the garnish. If you ever design a cache, steal the ordering.
Third, an honest caveat. The 890 bytes are the logical size. Some SGLang paths still reserve space in a FlashMLA-compatible layout (FlashMLA is the production attention-kernel library whose per-entry layout those paths were built around), so the bytes actually set aside per token differ. Kernel compatibility can quietly re-inflate a budget you had already spent on paper.
New piece three: an index, so the model does not re-read everything. Even a small cache grows with the prompt, and V4.1 refuses to attend over all of it. Its answer is an index added to the bank: before reading, a cheap set of search keys lets the model select the top-512 entries worth actually reading, so the expensive part of attention becomes a constant instead of a function of context length. That constant is the bill round 14's indexer kernels collect on. For the full machinery (two budgets, the hierarchy, and a worked 100K-token pass), open the aside:
The trick has two parts, and both fit in one picture. Attention needs two different jobs done per step: a cheap search (which past entries are worth reading?) and the actual read (mix their contents). V4.1 gives them different tools and different budgets. The search runs on separate, smaller keys: indexer keys are 128-dimensional, where the content vectors they point at are 512-channel, so scoring every position is affordable. The read then touches only the winners: the layer's query attends the top-512 selected entries, and its cost stops growing with context altogether.
Worked once, with numbers, at a context of 100,000 tokens. The decoder's very first layer is the only one that scans all 100,000 positions; while doing so it also picks the strongest 2,048 blocks of 8 positions, a shared candidate pool of 16,384, and hands it to everyone downstream. The heads of the later groups (the Reindex layers, 0-indexed 24, 28, 32, 36) score only those 16,384, about 16% of the context, and pick their own 512. The layers between them (Reuse) score nothing and read their group head's selection as-is. Total for twenty decoder layers: one full scan, four pool-only rescans, fifteen rounds of pure reading. Under V4's arrangement, every layer would have scanned. And the selection itself is per group, so one bank honestly serves twenty different readers.
Two more moves make this cheaper than its V4 ancestor, on top of the hierarchy: the index's keys are projected straight out of the main KV entries (in V4, the indexer had its own separate compression path from the hidden states), and CSA2 drops the overlap and the absolute positional embedding that CSA used while writing entries. The hierarchy is training-aware too: the candidate restriction is applied in post-training as well, so the heads learn under the same search domain they will meet at inference. The index cards themselves live in the same FP4 cache: they are the 68-byte K of the receipts.
New piece four: a phrasebook instead of more thinking. The extra 196B parameters are not more layers; they are Engram, a lookup memory. (The name is borrowed from neuroscience: an engram is the physical trace a memory leaves.) Concretely: two modules at layers 1 and 14 (zero-indexed) hash the last 2, 3 and 4 tokens separately, with 8 hash heads into tables of roughly 16M entries sized by distinct primes, and a context-aware gate decides how much of the retrieved embedding to mix in. Almost no FLOPs; one more piece of plumbing on the per-step path. The charming part for systems: a token's hash is known the moment the token exists, so its embeddings are deterministic addresses, prefetchable from host memory over background RDMA while earlier blocks compute. You will meet the gate again as round 10's fused Engram gate.
Step back for a second. These four pieces do different jobs, but they rhyme. Each is a small piece of math wrapped in fresh plumbing (indexing, normalization, fusion), and each sits on the path every decode step must walk, 40 layers deep. Small math on a hot path is exactly what a kernel team gets paid to chase. Hold that thought: every bar in the next section is one of these pieces, repaid.
Follow one token through the whole machine. Enough architecture; let a concrete prompt do the walking. Say it is three words: The sky is, and the model will add two more: blue, then a period.
Prefill first. The three words become three ids, then three hidden vectors of width \(d=5{,}120\), and flow through the first twenty layers. Layers 0 and 1 use only the 128-token window, which right now is the whole prompt. In the CSA2 groups, the three writers (layers 2, 8 and 14) each fuse adjacent token pairs into FP4 entries as they pass: The sky is one entry, is waits raggedly for its partner, so each writer holds about two entries when prefill ends. Between attention blocks, the MoE runs (8B active per token on this side), mHC keeps mixing its four streams, and Engram's tables answer three hash lookups per token for the short pairs and triples they see. At the boundary, the first decoder layer projects its full-resolution share from the encoder's output: three entries for our three tokens. Then the vocabulary head scores, and the model says its first word: blue. And notice what does not exist yet: any K or V for "blue". It was never an input, so the bank knows nothing about it. That changes only when it is fed back in.
One decode round now. "blue" re-enters at layer 0 and pays all forty layers, as the twist warned. At the encoder writers it completes the pair is blue: one fresh pooled entry each. At the first decoder layer it adds one full-resolution entry. Every decoder layer then asks its own question, reads the shared bank (for now, trivially all of it), attends its rebuilt window, and its Q dies on the spot. The logits come out as the period, and the round closes. Nothing about the question is stored; the window is still capped at 128; the bank alone has grown, by roughly one half-entry per encoder writer and one entry at the decoder writer.
The bookkeeping counts out with Rule 3's receipts (356 bytes the entry):
| after | entries in the global bank | bank size (≈) |
|---|---|---|
| prefill, 3 tokens | ≈ 2.5 × 3 ≈ 7–8 | ≈ 2.8 KB |
| decode step: "blue" | ≈ 10 | + ~890 B ≈ 3.7 KB |
| decode step: "." | ≈ 12.5 | + ~890 B ≈ 4.6 KB |
| a 4,096-token prompt | 10,240 | ≈ 3.6 MB |
Short prompts run above the 890-byte rate, because pooling only writes an entry once a pair completes; every further token settles onto the asymptote. And now the quiet pattern of the whole journey is visible: the query never accumulates, the window never outgrows 128 tokens, so the only thing in the machine that still grows with time is the bank, at about 890 bytes a step. Long-context cost is exactly that slope, and nothing else.
You need nothing new. Lectures 1, 6 and 7 are enough.
Three hiding places. One: a day-0 step is nowhere near its bandwidth floor: bytes wasted on fallback paths and extra round trips, tiny kernels idling in serial bubbles. Rounds 1 to 10 live there. Two: the four GPUs are not equally busy, and a tensor-parallel group (weights split across the four GPUs, Lecture 6) moves at its slowest member's speed; round 16's home. Three: change the unit of accounting: one step that drafts, verifies and accepts several tokens spreads its per-step cost over all of them: DSpark's job, rounds 11 to 15. Keep your list. Any bar you did not predict is the lesson to take home.
Why does decode care? At BS=1 a decode step reads roughly \(2\times(\text{active params})\) bytes of weight and the whole KV state, and does almost no arithmetic per byte. Do the arithmetic with the table's output row: about 16 billion active parameters per generated token, at about 2 bytes each, is on the order of 32 GB of weight traffic per generated token, shared across the four GPUs. That is a bandwidth floor no kernel rewrite can duck under, Lecture 1's roofline applied to this model. Lectures 1 and 7 told you what that regime rewards: small caches, short kernels, no gaps. V4.1 Flash hands you the first. The other two you must build. And once the step approaches that floor, exactly one lever remains: get more tokens per step. You can almost hear DSpark warming up.
Here is the whole journey at once. Rounds 1–10 are cumulative plain-decode measurements; rounds 11–16 switch to the final workload (random 4,096-token input, 1,024 output tokens, DSpark on with a simulated accept length of 5.5) and stay on it, so 11→16 are directly comparable. Hover or tap a round. Names in the tooltips you have not met yet (mHC, WO-A, split-K) are unpacked section by section below; read heights on the first pass.
Hover or click a round.
plain decode ·
DSpark · random 4K/1K · simulated accept length 5.5 ·
final: MoE TP4 + padding
Two honest caveats before we ascribe credit. First, the rounds are measurement checkpoints, not commits: several plain-decode steps landed bundled into the day-0 bring-up, and the post reports the checkpoints, not per-round diffs. Second, 35 → 203 (plain) and 203 → 874 (DSpark) answer different questions. The first asks "how fast can one token-step be?"; the second asks "how many draft tokens can verify make free?". And the DSpark table's own "speculation off" entry on the final workload is 223.5 t/s, so speculation itself contributes roughly a 3.9× of the final number.
Round 1→2: 35.2 → 117.8 tokens/s. The single largest step of the whole journey is not a
kernel anyone wrote. Some dense weights ship in FP8 (8-bit floating-point numbers plus one scale
shared per small block), but their scale layout (block_fp8) did not match what the
serving backend expected, so those matrix multiplies (the GEMMs) silently took a slow fallback
path. The fix: rearrange the scale layout once, when the weights are loaded
(block_fp8 → mxfp8 conversion in the quantization layer), and the same GEMMs go
straight into the Blackwell MXFP8 kernel, the same 8-bit values with their scales grouped the way
the hardware wants to read them.
Think: what does the fallback path pay per decode step that the native kernel doesn't?
The mismatched scale layout forces the kernel to either fall back to a much less hardware-specific implementation (no MXF8 tensor-core path, and poor use of the GPU's compute cores, the SMs, because the matrix has only one token row: M=1), or to dequantize/requantize on the critical path, which means extra memory trips per step on a workload that is pure memory traffic. At BS=1 a GEMM is bandwidth-bound; anything that multiplies its bytes or serializes its reads multiplies its time. Doing the rearrangement at load moves the cost off the per-step path entirely: one-time work, permanent free lunch.
Rounds 3, 6, 7: 117.8 → 133.5 → (146.5) → 148.4 → 152.1. Token-by-token decode runs a long chain of short kernels. Between RoPE (the rotary position transform from Rule 3), FP4 quantization, the compressor pooling, RMSNorm (the cheap root-mean-square normalization run between blocks) and the cache write, each op reads its neighbor's output from HBM (the GPU's main memory), launches, and writes its own. Fusing adjacent ops saves a launch and a round trip through memory, every time:
| Fusion | Implementation | BS=1 tokens/s |
|---|---|---|
| RoPE + FP4 | rotation, quantization and dequantization in one kernel | 117.8 → 133.5 |
| C2 compressor | the second compressor stage: normalization, pooling and the state write for adjacent tokens | 148.4 → 152.1 |
| WO-A, norms, Engram gate | WO-A is the attention output projection; as GEMV (matrix-vector product: the M=1 shape) for single-row projections; fused kernels for the small ops | 186.6 → 203.3 |
In between, a systems habit rather than a kernel: the team stopped rebuilding per-step request indices and scratch buffers in every layer and shared them across layers. That alone was 146.5 → 148.4. Once the fast paths were validated, they were turned on by default. Individual deltas look small; the compound effect is what carries 118 to 203. For calibration, the released stack reports that an entire Reuse-Mode layer executes with only 11 fused kernels during decode. That is the kind of budget a well-kept chain should fit into.
Rounds 4–5: 133.5 → 141.1 → 146.5. V4.1 carries mHC, a connector that keeps four residual streams instead of one. Every attention and MoE sublayer must compute mixing coefficients for the four streams and normalize them with Sinkhorn, a 20-iteration balancing routine that alternately rescales rows and columns until both sum to one, so no stream silently starves or dominates. Per call this is tiny, but it runs twice per layer (attention + MoE), and at small batch "tiny, twice, forty layers" is a real bill. The report's accounting fits on a napkin. Write \(n\) for the streams and \(d\) for the hidden width: here \(n=4\) and \(d=5{,}120\). The naive three-kernel version moves \((4n+4)d\) activations per block; the ideal residual map needs exactly \((n+1)d\) read and \((n+1)d\) written. Plug in the numbers: \(20d \approx 102{,}000\) values per block per token versus \(10d \approx 51{,}000\). The fixes were embarrassingly classical: pick tile sizes based on the number of input rows, fold the normalization weights into the projections offline, then fuse the statistics reduction with the Sinkhorn normalization so the coefficients are computed without an extra kernel round trip. When nothing is expensive, count the kernels, not the FLOPs.
Rounds 8–10 take off: 152 → 186.6. The largest plain-decode gain after the FP8 fix comes from a dependency observation, not a kernel. In V4.1's single-pass mHC, the pre-mix uses coefficients produced by the previous sublayer (the report shifts the input-mixing coefficients one block back for exactly this reason, landing at the ideal \((2n+2)d\) activation traffic), so the current attention or MoE does not need this sublayer's statistics yet. Two things that don't depend on each other can run at the same time: put the fused pre-mix + RMSNorm / statistics + Sinkhorn on side streams, attention or MoE on the main stream, and join them at the post-mix.
With the same fusion-and-overlap applied to the compressor and the indexer chains, plain decode reaches 186.6, and the WO-A GEMV row of the fusion table above lands it at 203.3. At this point the M=1 fast paths are validated and become defaults. Day-0 plain decode has grown 5.8×.
Round 11: DSpark arrives, at 558.24 tokens/s before optimization. DSpark ships with the official checkpoint: three lightweight draft blocks that take hidden states from the last few layers of the main model, compute base logits for five draft positions in parallel, and resolve the dependencies between draft tokens with a Markov head before handing the block to the target model for batched verification (Lecture 8's trick, factory-installed). A confidence head and a scheduler adapt the verification length to system load in production; in these runs, the block size stays fixed at 5 and acceptance is simulated at a target of 5.5 accepted tokens per step.
And that number, 6, breaks an assumption the entire left half of the ladder rested on: plain decode handles one row per request, and every fast path so far was shaped for \(M=1\): the GEMVs, the tile choices, the fused little ops. Verify at block size 5 with the anchor handles up to six rows. Those are still nowhere near a real GEMM shape, but they are not \(M=1\) either. All the small-batch paths had to be reworked. That is the subject of rounds 12–15.
Four groups of optimizations, each a PR, each answering "what is the verifier waiting on now?". The numbers are cumulative, measured on the final random 4K/1K workload.
| Stage | What changed | BS=1 tokens/s |
|---|---|---|
| DSpark, before optimization | draft blocks + Markov head work; M=1 paths un-reworked | 558.24 |
| + verify / MoE fusion & overlap | mHC overlap wired into verify and draft; WO-A writes directly the layout the next stage reads; candidate mask fuses the valid-length check with candidate handling; router emits the layout the experts need; MXFP8 input quantization overlaps routing on a side stream; expert-weighted reduction + shared-expert add + all-reduce fused into one kernel | 718.75 |
| + small-batch projections / mHC | WO-A uses split-K so more thread blocks work at once; mHC fuses the four-stream mixing with RMSNorm; the draft's KV projections reuse the MXFP8 weights and scales instead of the old FP8 path | 761.71 |
| + indexer / Q-RoPE / WO-A quant | post-Top-K score check + invalid-position filtering + KV page-address translation as one step; chosen candidate blocks expand straight into a token mask; Q's RoPE merged into the attention buffer write; WO-A's split-K reduction emits the following MXFP8 quantization itself | 802.38 |
| + C2 verify compression fusion | verify positions are contiguous, so within a request the previous row is read directly; only the first row needs the ring buffer; pair pooling + RMSNorm + RoPE + quantization + main KV write fuse into a single kernel, reused for index-K | 853.49 |
Notice the shape of the fixes: it is the same four skills from plain decode, applied to a batch of 6. Overlap (mHC, input-quant on a side stream), fusion (masks, reductions, quantization into the preceding kernel's epilogue), dispatch to the right kernel (MXFP8 instead of legacy FP8), and small-batch geometry (split-K for \(M \le 6\) so the SMs are not idle while one block does all the work). New model, new setting; the list of suspects was already memorized. That is what "experience with kernels" actually consists of.
853.49 → 873.63 tokens/s (+2.22% over the same-round EP4 measurement of 854.64). The last step is not a kernel at all; it is a parallel-layout decision (Lecture 6). At BS=1, expert parallelism means each GPU owns a disjoint set of experts: if a routing's token lands unevenly, one GPU has more work and the other three wait at the all-reduce. Tensor parallelism makes every GPU compute a different slice of the same experts, so uneven expert load spreads across all four ranks instead of piling onto one.
The catch is geometry. Each expert's intermediate width is 2,304 channels; split by TP4 that leaves 576 per GPU, and the kernel wants alignment to nice tile sizes. So the weights are padded to 640 per GPU at load time (zeros on the weights, ones on the scales so padded channels compute to zero). The machine does strictly more compute per step... and finishes sooner, because at BS=1 the real bottleneck was never FLOPs; it was the slowest rank (a rank is one GPU of the group). In the trace, the median gap between ranks arriving at the all-reduce barrier (the collective where the four GPUs exchange partial sums) dropped from 11.14 µs to 3.40 µs.
Remember Lecture 6: what does a TP group finish at the speed of?
A tensor-parallel layer finishes when its slowest rank finishes; every rank-and-file imbalance is paid four times over at each barrier. Under EP4, an unlucky routing concentrates work on one GPU and the barrier gap (11.14 µs median here) is dead time on three cards. Under TP4 all GPUs see the same token and compute matched slices, so the barrier gap collapses (3.40 µs). At BS=1 the per-rank GEMMs are far from compute-bound, so the 11% extra (padded) arithmetic is absorbed by idle tensor-core cycles that were going to waste anyway; what you buy back is balance. Tiles, not totals, decide this regime.
A 25× claim deserves a precise protocol, and the post gives one. Rounds 11–16 use:
4,096 random input tokens (ids drawn uniformly with fixed seed 42, special tokens excluded,
no chat template; every configuration reuses the very same ids), a fixed 1,024 output
tokens with temperature=0 and ignore_eos=True, and a simulated
accept length targeting 5.5 (match-expected accepts 5 or 6 tokens per round;
measured values like 5.505 and 5.520 are kept as-is since finite rounds and last-step truncation
shift the expectation slightly). Acceptance is simulated, so the text is not used to judge
quality, and simulation mode disables the in-graph acceptance path.
Throughput counts tokens added after the first streamed event divided by first-to-last event
time, excluding full prefill. Each server launch is measured over 6 runs (one warm-up
discarded, cache cleared between runs); the contested configurations were launched twice
independently with the median over 12 runs, alternating TP/EP launches for the final comparison,
profiler off. The serving command behind 873.63, from the root of the pinned checkout
(BBuf/sglang@835c3909, official checkpoint at revision dba1be0a,
PyTorch 2.13.0+cu130, FlashInfer 0.6.18, Triton 3.7.1):
SGLANG_RAGGED_VERIFY_MODE=static \
SGLANG_SIMULATE_ACC_LEN=5.5 SGLANG_SIMULATE_ACC_METHOD=match-expected \
CUDA_VISIBLE_DEVICES=0,1,2,3 PYTHONPATH="$PWD/python" \
python -m sglang.launch_server \
--model-path "$MODEL_PATH" \
--served-model-name deepseek-ai/DeepSeek-V4.1-Flash \
--tp 4 --ep-size 1 --trust-remote-code \
--moe-a2a-backend none --moe-runner-backend flashinfer_mxfp4 \
--mem-fraction-static 0.80 --max-total-tokens 33554432 \
--chunked-prefill-size 4096 \
--cuda-graph-bs-decode 1 2 4 8 16 32 64 \
--max-running-requests 128 \
--speculative-algorithm DSPARK --speculative-dspark-block-size 5 \
--skip-server-warmup --reasoning-parser deepseek-v41 \
--random-seed 42 --decode-log-interval 10 \
--host 127.0.0.1 --port 30021
--tp 4 --ep-size 1 means both attention and MoE use TP4, and this code applies the
intermediate-dimension padding when it loads the weights. For the EP4 comparison: change
--ep-size 1 to --ep-size 4, everything else identical. And the same-run,
same-input summary the post ends with:
| Workload and metric | DSpark off · EP4 | DSpark on · EP4 | DSpark on · TP4 + padding |
|---|---|---|---|
| BS=1, output speed (tokens/s) | 223.50 | 853.49 | 873.63 |
| Measured accept length (target 5.5) | none | 5.520 | 5.505 |
The rounds are measurements; this is the code that produced them, as it landed upstream in
SGLang between September 9 and 12, 2026. (The reproduction pins BBuf/sglang@835c3909,
the author's fork; the numbered PRs below are the upstreamed commits.)
| Commit | Date | What it carries |
|---|---|---|
85f8105f5d "Add DeepSeek V4.1 Day 0 support" | Sep 9 | The model + the first kernel generation: model file, FP8/FP4 paged KV layouts, RoPE+FP4 fusion (fp4_rope), C1/C2 compressors, FP4 indexer, sparse Top-K v2, mHC kernels, Engram kernels, WO-A GEMV, small GEMMs, the block-FP8→MXFP8 load-time conversion, and the expert intermediate-dim padding used by round 16 |
69777c4d36 cookbook (#38802) | Sep 9 | Day-0 launch recipes |
9e57251377 FlashMLA bump | Sep 10 | Attention kernel fork at the V4.1 rebase head |
4b7331fb77 "remaining NVIDIA cells start" (#38861) | Sep 10 | The fast paths/configs turned on everywhere |
c36636b7da DSpark verify & MoE (#38879) | Sep 10 | Round 12 → 718.75: mHC overlap in verify/draft, candidate-mask fusion, MoE input-quant overlap, fused reduce+shared-expert+all-reduce |
759baff47f small DSpark batches (#38976) | Sep 11 | Rounds 13–14 → 761.71 / 802.38 ("BS1 803 tok/s" is literally in the title): WO-A split-K, mHC mixing+RMSNorm fusion, draft KV on MXFP8 |
0d5e663b8f verify compression, indexer & projections (#39068) | Sep 11 | Round 15 → 853.49: fused indexer post-processing, Q-RoPE into the buffer write, WO-A emitting MXFP8, the C2 verify kernel |
da64c5cbb8 raw-index TopK v2 (#39098) | Sep 11 | Indexer returns positions without materializing masks |
3a42fc5571 Engram decode history (#39138) | Sep 11 | N-gram history committed inside the hash kernel |
93ca76741c two-level candidate indexer (#38944) | Sep 12 | CSA2's hierarchical candidate filter on DeepGEMM's paged sparse MQA logits |
53ff45a3d6 Hopper decode opts for TP8 | Sep 12 | H200 tuning follow-up (not in the GB300 ladder) |
None of these results use the new kernels DeepSeek released with V4.1. Everything in the ladder is in-house work on top of the public stack.
When the next unfamiliar model lands on your desk and its day-0 build is "slow", this is the order the journey suggests. Each line matches a section above. Memorize the suspects, not the numbers.
A first working build optimizes for correctness, so its waste is the coarsest kind: whole subsystems running on fallback paths, extra copies of everything. Each later fix is spent fighting the remainder, and multiplicatively so: removing 30% of a step that used to be 70% of the runtime is worth less than removing it when it was 90%. Nothing deep, but it predicts where to look first on the next model (dispatch and fallback audits), and it explains why case studies bunch their big numbers early.
At M=6 (six token rows hitting weight matrices of size N×K), a plain GEMM launch has few tiles in the M direction: too few thread blocks (CTAs) to fill the SMs, the GPU's compute cores. Splitting the reduction dimension K across blocks creates the missing parallelism; a small in-kernel (or fused-epilogue) reduction merges partials. The cost is the extra partial-sum traffic and the reduction; that is why WO-A's split-K reduction was later fused with the following MXFP8 quantization: the merge happens once, in the epilogue, and no intermediate HBM tensor exists.
To pool adjacent token pairs you need the previous token of every position; for random positions that is a scatter read through a ring buffer. Verify rows within one request are consecutive in the batch, so "previous" is just the previous row: a fixed-offset read, free. Only the batch's first row per request still needs the ring. Anything that reorders or interleaves rows (ragged verify modes, cross-request packing) destroys the invariant; the fused kernel is guarded by exactly that precondition.
Spot the missing kernel. The post ends with the line: "None of these results use the new kernels DeepSeek released with V4.1." The ladder's last step was +2.2%. The report names the released set: FlashMLA's fused RoPE-attention-RoPE-cast pipeline, Mega-Gate / Mega-mHC / Mega-MoE in DeepGEMM, TileKernels, and DeepSelect's TopK. List what they cover that this journey fused by hand, and estimate whether plugging them in is a new round 17, or a rewrite of rounds 3–15 with a bigger baseline. There is no right answer; there is a lunch argument.
Thanks to the DeepSeek team for open-sourcing DeepSeek-V4.1, and to everyone on the SGLang and Miles teams and in the community who helped with the model support, the kernels, testing and review. Some of the kernels were developed with the KDA 0.5 framework, and we are grateful to Humanize and Kernel Design Agents for the tools and the workflow.
This lecture is retold from
the SGLang team's post "From 35 to 873 tokens/s: the kernel optimization journey" and the SGLang
repository history (the commits above), with the measurements as reported there: 4× GB300, BS=1,
attention TP4, random 4K/1K, simulated accept length 5.5. The architecture details in "The
Setting" (the Full/Reindex/Reuse mode layout, the 890-byte accounting, the FP4 format, Engram,
DSpark, and SWA Bounded Replay) follow the DeepSeek-V4.1-Flash technical report. The KV-cache
fundamentals (why Q is not cached, one shared bank under sparse selection, the shared K=V
representation) were cross-checked against Dongxi 东锡's public KV-cache explainer on X; its
five animations are redrawn as step-by-step figures in assets/kv-steps/.