Strata is an open-source local inference engine built on llama.cpp/ggml that runs the 125-billion-parameter Qwen3.8-Flash-Next MoE model on consumer gaming PCs with as little as 12GB VRAM, using tiered expert caching (GPU→RAM→SSD) and guess-and-check speculative decoding to achieve 53-94 tokens/second.
As an Amazon Associate, we earn from qualifying purchases.