OPEN SOURCE DEEP DIVE
Wan2GP: run frontier generative models on a 6 GB GPU
A one-stop local generation app for video, image, audio and speech models, built on MMGP: block swapping, VRAM preloading and seven memory profiles let models larger than your VRAM run, starting at 6 GB.
What it is
Wan2GP (deepbeepmeep/Wan2GP, 10,098 stars / 1,603 forks, Python) describes itself as a fast AI video generator for the GPU poor. It is a one-stop local generation app that puts video, image, audio and speech models behind a single Gradio interface so they run on consumer graphics cards. The name is a pun: 2GP means GPU Poor, the same models that need an A100 elsewhere start at 6 GB of VRAM here.
But the part worth pulling out is not the model list, and not the 14,389-line main program. It is the memory management layer underneath, MMGP. That is a separately reusable library, now living in the same repository, built specifically to handle "the model is larger than the VRAM", and it is the technical floor that lets this app carry dozens of models at once.
The numbers
| Field | Value |
|---|---|
| Repo | deepbeepmeep/Wan2GP (site wangp.ai) |
| Stars / Forks / Watchers | 10,098 / 1,603 / 111 |
| Language / Licence | Python / WanGP Community License 2.0 (custom, not OSI open source, see below) |
| Created / last commit | 2025-02-27 / 2026-10-06, still shipping frequently |
| Tracked files | 2,807 (core code, per-model adapters, several hundred defaults/*.json model presets) |
| Minimum VRAM | 6 GB for select models; 15 seconds of 1080p H3 video reportedly went from 25 GB to 11 GB |
| Hardware | NVIDIA GTX 10XX / RTX 20XX and newer (older cards included); AMD RDNA 4 / 3 / 3.5 / 2 |
| Quantisation formats | int8, fp8, GGUF, NV FP4, Nunchaku, plus INT8 ConvRot |
| Install | Windows 10/11 and Linux; conda plus pip install torch==2.10.0 (RTX 20XX+) or 2.7.1 (GTX 10XX); a Dockerfile is included |
Architecture: two layers, not one
Reading Wan2GP as "a video generator" misreads its structure. It is two layers, and the lower one can be taken out on its own:
| Layer | What it is | What it does |
|---|---|---|
| MMGP ( mmgp/, v4.0.0) | A general memory management layer, "Memory Management for the GPU Poor", originally its own project, now merged into this repo | Block swapping, VRAM preloading, reserved RAM, pluggable quantisation, merge-free LoRA, its own VRAM allocator |
| WanGP ( wgp.py and friends) | The application on top: model adapters, Gradio UI, generation queue, galleries, plugins, the Deepy agent | One multimodal entry point, production workflow, API |
MMGP states plainly that it replaces accelerate's offloading, because the latter handles "a model loaded and unloaded several times in one pipeline, a VAE say" badly. That sentence names the project's real concern: on consumer hardware the bottleneck for video models is not compute, it is the waste from shuttling the same model across the VRAM boundary over and over.
Most of the 300k lines of Python are per-model adaptation layers under models/, but the genuinely original scheduling code is concentrated: wgp.py at 14,389 lines is the main program, the memory core mmgp/offload.py is 5,225 lines, and mmgp/allocator/ is a C++ extension shipping a Linux .so and a Windows .dll. The presence of MMGP's own sub-project README inside the repo is the tell: it was designed as a reusable component, not only as a bundled client.
How MMGP turns "bigger than VRAM" into a choice
Four mechanisms, and the combination is the actual technical content:
- Block swapping. A model larger than its VRAM budget is processed block by block. Blocks are copied to the GPU ahead of their use while the previous ones compute, in an order learned during the first step (including from one tower of blocks to the next and from one step to the next), landing in fixed VRAM slots carved out of that model's budget. This is what lets a 6 GB card run a model that does not fit.
- VRAM preloading. The part of the budget not needed by blocks in transit keeps blocks resident in VRAM, so fewer blocks move per step. Version 4.0 chooses this together with reserved RAM so the two do not fight each other.
- Reserved RAM. Kept in pinned memory (copied in parallel, then locked), transfers run up to 2 times faster. That reserved block is shared between models by priority: a model that does not fit gets the blocks it needs at each step first, spread evenly, and the leftovers go through a small staging ring that stays nearly as fast as pinned models when the GPU is busy. There is also smartPinning for the models left out, and Dynamic VRAM Preload which fills spare VRAM with weights and frees it as needed.
- Pluggable quantisation. Handlers convert other checkpoint formats to
quantoas they load (scaled FP8 is built in); WanGP itself adds GGUF, NV FP4, Nunchaku and INT8 ConvRot.
Two more details that are easy to miss and very practical: LoRA, DoRA, LoKr and diff adapters are applied during the forward pass without merging, so they work on quantised models too, and the multiplier can change between denoising steps. And the MMGP Optimized VRAM Allocator replaces PyTorch's VRAM allocator outright on Windows and Linux, recycling unused VRAM more efficiently and saving several GB of peak on long videos and large images.
Getting the most out of it needs two upstream settings: Sage2/2+ Attention, plus an MMGP Optimized VRAM Allocator option and Smart Memory Pinning under Config → RAM/VRAM Management. The README also quantifies the trade-off: setting Attention Head Split to Medium saves another 20% VRAM for up to a 10% speed penalty.
Profiles: the memory trade-off as a product option
This is where the craft shows. Most comparable projects ship one "low VRAM mode" switch; this one ships seven, each a different point on the trade-off curve, each annotated with the situation it suits. Pick one in the UI and the memory layer rebuilds its scheduling plan around it.
| Profile | Strategy | Good for |
|---|---|---|
| 1 HighRAM_HighVRAM | Each model loaded whole in VRAM, all models in reserved RAM | Fastest generation and switching; needs the most RAM and VRAM |
| 2 HighRAM_LowVRAM | All models in reserved RAM, sent part by part | Runs models larger than VRAM, leaves VRAM for long video, fast switching; needs lots of RAM |
| 3 LowRAM_HighVRAM | Each model whole in VRAM, only the main models in reserved RAM | Fast with less RAM; VRAM must fit the whole model |
| 3+ VeryLowRAM_HighVRAM | Profile 3 with no reserved RAM at all | Recommended for audio: models are small enough to fit in VRAM, where their language model runs much faster |
| 4 LowRAM_LowVRAM | Only main models in reserved RAM, sent part by part | Recommended for most PCs: runs oversized models, keeps most VRAM free, small speed cost |
| 4+ LowRAM_LowVRAM+ | Profile 4 sending one part at a time | Saves roughly 1 GB more VRAM, slightly slower |
| 5 VerylowRAM_LowVRAM | Almost no reserved RAM, everything sent part by part | Fail-safe for machines short of both; noticeably slower on short jobs like images |
Two details in the config panel are worth quoting. Profile 4 is labelled the recommendation for most PCs, and Profile 3+ is the audio default because audio models are small enough to live entirely in VRAM where their language model runs faster and can use the faster CUDA Graph or vLLM engines. The panel also states the cost plainly, that reserved RAM cannot be used by other programs, so only part of your RAM is set aside. Naming the downside is better than letting users discover it.
Model coverage
It trains nothing. Its value is making other people's models runnable. The current README spans three modalities:
| Modality | Models |
|---|---|
| Video | Wan 2.1 / 2.2 and derivatives, MiniMax H3, LTX-2 / 2.3 / 2.5, Hunyuan Video 1 / 1.5, LongCat, Kandinsky, LTXV, MagiHuman |
| Image | Krea 2, Qwen Image, Z-Image, Flux 1 / 2 (Klein, Chroma), SenseNova, Ideogram 4, HiDream |
| Audio / TTS | Qwen3 TTS, MiniMax H3 Voice Clone, Ace Step 1 / 2 / XL, Omnivoice, Index TTS2 / 2.5, KugelAudio, HeartMula, Chatterbox, Minimax Music, Stable Audio 3 |
Around those models an engineering layer has grown: LoRA customisation (including reusing LoRAs stored in another app), finetunes (your own, or from Hugging Face and CivitAI), many quantised checkpoint formats, an architecture-aware downloader that fetches files suited to your hardware, a generation queue you can walk away from, a headless mode for batch jobs, and the WanGP API. On the output side there are galleries, reusable settings templates, a per-model prompt enhancer, and a set of pre/post tools: mask editor, background remover, pose / depth / flow extractors, speaker diarisation, background noise and song removal, RIFE and FlashVSR temporal and spatial upsampling, MMAudio soundtracks, SeedVC voice replacement, and remuxing a video with any soundtrack.
More than a model shell: Deepy and the production path
Two decisions move it from a toy into a production tool.
First, Deepy, a low-VRAM media agent that can work offline. It orchestrates the tedious parts: transcription, splitting, colour-frame extraction, masking named objects, replacing or mixing audio tracks while keeping selectable tracks and subtitles without re-encoding the video. There are two tiers. Deepy Zero is light and fast, for single tasks: generate one asset, edit selected media, extract a clip, produce a transcript. Deepy Prime is for work that needs planning or several connected actions: comparing compatible models, combining multiple media assets, inspecting intermediate results, managing project files, working with external MCP services. Prime needs a local Qwen3.8 VL (9B or 27B) or a configured remote LLM. The docs are candid about the limit: Deepy can make mistakes, so verify important results.
Second, the API and plugin surface. shared/api.py is an in-process wrapper whose stated goals are letting third-party code call WanGP directly, keeping the last loaded model alive across requests, delivering structured progress updates, and still capturing the stdout/stderr that would normally go to the console. The same API is usable from a WanGP plugin and works with WanGP Web Queue to process jobs submitted by a third-party app.
Two more capabilities are easy to overlook. There is a Comfy Kitchen kernel chain that auto-selects Comfy Kitchen CUDA (HIP on AMD), then Triton, then PyTorch, including fused INT8 linear math and INT8 ConvRot decode kernels, and reports the backend actually chosen at startup. And there is optional DLSS 5 neural rendering, usable as a native-resolution refiner or as a spatial and temporal upsampler; the project is explicit that it pulls in closed-source third-party binaries with separate licences and security implications, installed by a dedicated checksum-verified script.
Licence: the section you must actually read
GitHub reports the licence as NOASSERTION, because this is a custom WanGP Community License 2.0 (455 lines), not an OSI-approved open source licence. If commercial use is on the table, this section matters more than any feature list:
- "Free Use" is broad: personal, hobby, research, educational, evaluation, internal company use, studio, agency and client work, implementation and support work, private deployment, free redistribution.
- Free Use explicitly covers batch mode, queue mode, headless usage, local-service mode, scripted, automated and API mode — so running WanGP as an internal production pipeline needs no extra licence.
- Restricted Commercialization is five things: selling, sublicensing, renting, leasing, licensing for a fee or otherwise monetising access; offering the software through a paid, metered, sponsored, ad-supported, subscription, hosted, managed, API, SaaS, white-label, OEM, embedded, marketplace or platform product; distributing it as a material feature of a paid product; letting third parties invoke or consume it through headless usage, an API, an integration, a local service, a plugin, an IPC bridge or a remote endpoint in exchange for consideration; and charging a separate fee to unlock, activate, host, enable, bundle or provide access.
- Explicitly not restricted: internal company or business use; creating outputs for yourself or for clients; charging reasonable service fees for installation, customisation, consulting, support, training or integration labour, as long as no separate fee is charged for access to the software itself; and Direct Output Sales under section 6.
- Section 3.3 adds an explicit clarification: merely being a company, studio, agency or revenue-generating business does not require a separate reseller or commercial licence, as long as the software is not sold or made available to third parties in the ways section 5 describes.
The README makes the same point in prose: running it locally is always free, and the official project will never ask for a licence fee, subscription or donation to run WanGP on your own computer. It also warns to use only the official GitHub repository or wangp.ai / wan2gp.ai, and that WanGP is not affiliated with any third-party service using the WanGP/Wan2GP names.
Size and activity
2,807 files, several hundred defaults/*.json model presets organised per model and precision, and documentation covering thirty-odd topics (installation, models, LoRAs, plugins, API, CLI, Deepy, DLSS 5, VACE, upsamplers, troubleshooting). The release cadence is tight: between 2026-09-24 and 10-06 alone there were seven versions including Qwen Image 2.1, a community release, LTX-2.5 VFX tools and v17.00. The headline numbers for v17.00, which merged MMGP v4 into the repository, are that 15 seconds of 1080p H3 video went from 25 GB of VRAM to 11 GB (5–6 GB at 480p), Profile 4 up to 25% faster and up to 50% less VRAM, and the fail-safe Profile 5 up to 50% faster. The author also notes, in passing, that the project just hit 10,000 stars, with a stated goal of 100,000.
Who it is for, and who it is not
Good fit: anyone who wants the latest video, image and audio models running on their own GPU; people whose VRAM is short but who are willing to deal with quantisation and profiles; studios producing video locally who need a queue and headless mode; and anyone who wants an API to build generation into their own product.
Poor fit: anyone planning to wrap the software itself into a paid or hosted service, which is Restricted Commercialization; anyone who wants zero-setup operation and is unwilling to touch a Python environment, since the project explicitly rejects PyTorch 2.8.0 and 2.9.0 and RTX 20XX versus GTX 10XX need different Python and CUDA combinations; and compliance teams that need an OSI-approved licence, because this is source-available, not open source.
One more habit the author keeps insisting on: get it only from the official repository or website. Third-party repacks and renames are a known problem for projects with an install base this size.
Bottom line
The real contribution is not how many models it collects. It is that Wan2GP turned "the model does not fit in VRAM" into an engineering solution with seven selectable profiles, a reusable core library, and a quantisation format router. MMGP now stands as a library in its own right and is reused elsewhere, which suggests the abstraction holds up. For anyone running open generative models locally, it amounts to a portable VRAM budget manager — which happens to be the single biggest bottleneck every open generation model shares on consumer hardware.
SOURCE LINKS