
X-NavDP: Generalizing Navigation Diffusion Policy to Novel Behavior and Embodiments with Group Q-score Reweighted Matching
Proposes GQRM, a data-efficient diffusion RL post-training framework with self-bootstrapped exploration and group Q-score normalization for cross-embodiment visual navigation. Improves success rate from 61.20% to 84.28% in simulation and 10% to 65% in real-world hard cases.
Tianyu Yang, Yiming Zeng, Wenzhe CaiJul 30, 2026
NavDPVLANavigationJul 30, 2026