OPEN SOURCE DEEP DIVE
SGLang: the serving framework that made prefix reuse a radix tree, and a rollout backend for frontier RL post-training
A high-performance serving framework hosted by LMSYS: RadixAttention prefix caching, a zero-overhead CPU scheduler, prefill-decode disaggregation, DFlash and Spec V2 speculative decoding, compressed finite state machines for structured output, and large-scale expert parallelism (96 H100 GPUs; 3.8x prefill and 4.8x decode on GB200 NVL72 part II). Its day-0 ledger covers Kimi K3, DeepSeek-V4, GLM5.2 NVFP4, Nemotron 3 and MiniMax M2, and AReaL, Miles, slime, Tunix and verl all use it as an RL rollout backend.
RadixAttention: making prefix reuse a first-class engine structure
SGLang is hosted by LMSYS, a non-profit open-source organisation, and is a high-performance serving framework for LLMs and multimodal models. Its signature is RadixAttention: a radix tree organises the KV cache already resident in VRAM, so "several requests share the same prefix" stops being an accidental cache hit and becomes the engine's core data structure. Anything that shares a system prompt, few-shot examples, multi-turn history or agent tool definitions recomputes only the delta instead of the whole prefix. The January 2024 launch blog claimed up to 5x faster inference.
This route is complementary to vLLM's paged memory rather than opposed to it: paging decides how blocks are placed, the radix tree decides which blocks can be reused and in what order the longest common prefix is matched. SGLang is explicit in its acknowledgements that it learned from and reused design and code from vLLM, FlashInfer, Guidance, Outlines, LightLLM and LMQL - projects at this layer copy each other constantly, and that is normal rather than scandalous.
The rest of the runtime: a zero-overhead scheduler and large-scale expert parallelism
The README unpacks "Fast Runtime" into a flat list: RadixAttention prefix caching, a zero-overhead CPU scheduler, prefill-decode disaggregation, speculative decoding, continuous batching, paged attention, tensor/pipeline/expert/data parallelism, structured outputs, chunked prefill, quantisation (FP4/FP8/INT4/AWQ/GPTQ) and multi-LoRA batching. Two of these represent its main engineering investment beyond the pack:
| Capability | Problem it solves | Published result |
|---|---|---|
| Zero-overhead CPU scheduler (v0.4, 2024-12) | Scheduling logic runs on the CPU and competes with GPU compute for wall time; the larger the batch and the faster the model, the more likely the CPU side becomes the bottleneck | v0.4 shipped it alongside a cache-aware load balancer and faster structured outputs |
| Large-scale expert parallelism + P/D disaggregation | MoE models carry more experts than one card holds, and prefill and decode want different optimal parallel strategies - mixed together they drag each other down | deployment blog on 96 H100 GPUs (2025-05); GB200 NVL72 part II reports 3.8x prefill and 4.8x decode throughput (2025-09) |
| Compressed finite state machine | Structured output such as JSON validates grammar token by token, and computing the mask becomes its own overhead | 2024-02 blog: 3x faster JSON decoding |
| DFlash and Spec V2 | Draft quality and acceptance rate in speculative decoding | the 2026-06 blog calls it "the next generation of speculative decoding"; DFlash is also listed by vLLM as one of its routes |
The day-0 record: the metric that actually measures this race
For a serving engine the most persuasive evidence is not a steady-state throughput figure but whether a new model runs on release day, and how quickly it reaches production quality. SGLang's News list is effectively a day-0 ledger:
| Date | Event |
|---|---|
| 2026-07 | Day-0 support for Kimi K3 together with Miles; RadixArk and Google bring the full SGLang feature set to TPUs; serving GLM5.2 NVFP4 agentic workloads reaches 500 TPS in two weeks |
| 2026-06 | DFlash and Spec V2; day-0 for Nemotron 3 Ultra / Super and Higgs Audio v3 TTS |
| 2026-04 | DeepSeek-V4 day 0: from fast inference to verified RL with Miles |
| 2026-02 | 25x inference performance unlocked on NVIDIA GB300 NVL72 |
| 2026-01 | SGLang Diffusion accelerates video and image generation |
| 2025-12 | Day-0 for MiMo-V2-Flash, Nemotron 3 Nano, Mistral Large 3, LLaDA 2.0 diffusion LLM and MiniMax M2 |
| 2025-10 | SGLang-Jax backend, running natively on TPU |
| 2025-09 | Day-0 for DeepSeek-V3.2 with sparse attention |
Pulling diffusion models (WAN, Qwen-Image) and TTS into the same runtime is what separates it from a pure LLM engine: SGLang is aiming to be the unified serving layer for generative models, not only for autoregressive decoding.
Its other identity: the RL rollout backend
The README devotes a section to "RL & Post-Training Backbone": SGLang is a proven rollout backend used to train several frontier models, with native RL integrations and adoption by post-training frameworks including AReaL, Miles, slime, Tunix and verl. This matters. In RL post-training, rollout throughput sets the wall-clock time of the whole run, and rollout in turn demands hot weight updates and per-step KV cache resets. Whether a serving engine can act as the sampler on the training side is the ticket into frontier labs, and SGLang's position there is firmer than its position on pure serving.
Scale and hardware
The project states it generates trillions of tokens per day in production across more than 400,000 GPUs. The adoption list spans xAI, NVIDIA, AMD, Intel, LinkedIn, Cursor, Oracle Cloud, Google Cloud, Microsoft Azure, AWS, Baseten, Baidu, AntGroup, Alibaba and Tencent, plus MIT, UCLA, the University of Washington, Stanford, UC Berkeley and Tsinghua. Hardware covers NVIDIA (GB200/B300/H100/A100/Spark/5090), AMD (MI355/MI300), Intel Xeon CPUs, Google TPUs and Ascend NPUs. In June 2025 it received the third batch of a16z's Open Source AI Grant.
The provenance of those numbers needs stating: 400,000 GPUs and "trillions of tokens a day" are the project's own figures with no third-party audit, so we record them as project claims. The day-0 ledger, which post-training frameworks adopt it, and LMSYS's governance status are verifiable facts, and those three carry the judgement.
One telling detail: SGLang offers coding-agent sponsorship to long-term active contributors - Cursor, Claude Code or OpenAI Codex - claimed by email with your most important commits or PRs. An open-source inference engine using AI coding subscriptions as contributor incentive is itself a footnote on the current state of this race.