Skip to content
←Back to Open Source

OPEN SOURCE DEEP DIVE

LLMInferenceOptimizationMoE

Strata: Run a 125B-Parameter MoE Model on Your Gaming PC

Strata is an open-source local inference engine built on llama.cpp/ggml that runs the 125-billion-parameter Qwen3.8-Flash-Next MoE model on consumer gaming PCs with as little as 12GB VRAM, using tiered expert caching (GPU→RAM→SSD) and guess-and-check speculative decoding to achieve 53-94 tokens/second.

Niko1221/Strata7.8kC++MIT4 min read

Overview

Strata is an open-source local inference engine that runs the 125-billion-parameter Qwen3.8-Flash-Next model on ordinary gaming PCs. Built on llama.cpp / ggml, it supports NVIDIA (RTX 20-50 series) and AMD (RX 6000-9000 series) GPUs with as little as 12 GB of VRAM. All inference happens locally — no data leaves your machine.

The Problem

A 125-billion-parameter MoE model normally requires server-grade hardware with hundreds of gigabytes of graphics memory. Strata's core innovation is distributing inference work across the entire PC — GPU, RAM, CPU, and SSD work together to make consumer hardware capable of running a hundred-billion-parameter model.

Core Architecture

Strata's technical design revolves around two key ideas:

MoE Expert Tiered Caching

Qwen3.8-Flash-Next is a MoE model with 24,576 experts, but each token activates only 10 of them. Strata exploits this property with tiered caching:

  • GPU VRAM: Caches the few thousand most frequently requested experts (hot experts)
  • System RAM: Holds all 24,576 experts
  • CPU: Processes GPU cache misses in parallel
  • SSD: Stores a large lookup table, loaded on demand

Think of it like a kitchen layout: ingredients used all the time stay on the counter (GPU), the rest waits in the pantry (RAM/SSD), fetched as needed.

Guess-and-Check Speculative Decoding

A small helper model guesses the next few tokens, and the large model verifies all guesses in a single forward pass. Correct guesses are accepted immediately — effectively generating multiple tokens per forward pass, yielding a 1.6-1.8x overall speedup.

Chunked Long-Context Reading

Long texts (such as 32K-token documents or code) are read in large chunks of up to 8,192 tokens, achieving over 1,000 tokens/second input processing speed.

Performance Benchmarks

Measured on two ordinary gaming PCs:

ConfigQuantGenerationPrompt Processing
RTX 5070 (12GB) + 64GB RAMQ2_094 tok/s2,650 tok/s
IQ2_XS79 tok/s2,090 tok/s
IQ3_XXS62 tok/s1,750 tok/s
IQ3_S53 tok/s1,620 tok/s
RX 9070 XT (16GB) + 47GB RAMQ2_060 tok/s1,160 tok/s
IQ2_XS52 tok/s1,110 tok/s

60 tokens/second is faster than human reading speed. Larger-VRAM cards (e.g., RTX 3090 24GB) are expected to reach 100-140 tokens/second.

Model Variants and Quantization Levels

Strata offers multiple model variants to fit different RAM capacities:

System RAMRecommendedNotes
32 GBCoderCode-specialized with half the experts removed; 91% of full model's SWE-bench Verified score
48 GBIQ2_XS or Q2_0Larger levels don't fit
64 GBIQ2_XS (recommended) / IQ3_XXS / IQ3_SAll levels fit; IQ3_S is best quality, slowest
96 GB+IQ3_S or Unsloth 4-bitRoom for the largest levels

Additional variants: Swift 1.5 (thinks shorter, answers sooner at similar quality) and Unsloth UD-Q4_K_XL (experimental, closest to full model but mostly read from SSD at 7-8.5 tok/s).

Installation and Usage

Installation is one-click: Windows users double-click START-HERE.bat, Linux users run ./setup.sh. The installer auto-detects the GPU, selects the right engine and model, downloads ~70 GB of model files (with resume support), and opens the browser at http://127.0.0.1:8080.

Also supports automated installation via AI coding assistants (Claude Code, Cursor, Codex, GitHub Copilot) or management through its MCP server.

External interfaces:

  • OpenAI-compatible API: http://127.0.0.1:8080/v1
  • Anthropic-compatible API: http://127.0.0.1:8080/v1/messages (Claude Code: ANTHROPIC_BASE_URL=http://127.0.0.1:8080)
  • Supports image input, thinking level control (off/low/medium/high), multi-GPU distribution, phone/remote access

Technical Position

Strata belongs to the LLM inference optimization domain. Its core contribution is tiering MoE expert caches across the full storage hierarchy of consumer hardware (GPU → RAM → SSD), combined with guess-and-check speculative decoding to achieve near-server-class inference speeds. It is not a new model — it is an inference engine that makes existing large models usable on ordinary PCs, dramatically lowering the hardware barrier to local hundred-billion-parameter model deployment.

License

Strata is open-source under the MIT License. The underlying model Qwen3.8-Flash-Next is developed by the Qwen team, with quantized versions by ISTA-DASLab, UkisAI, and Unsloth. The engine is built on llama.cpp / ggml. Individual components and models may carry their own licenses.

Related Projects