
DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation
DECOWAM adapts a frozen FastWAM video-action backbone to legged mobile manipulation via decoupled interfaces — an action-equivalent future bottleneck, adversarial base/arm factorization, and ego-motion-aware video conditioning — cutting Stage-2 trainable parameters 232x while leading real-robot deployment at 58.2% success.
Siyuan Ma, Boshi Zhang, Yutian ZhangAug 20, 2026
VLAWorld ModelsWorld-action modelsAug 20, 2026