用统一模型同时预测未来画面和动作,让机器人导航更准更可靠。
AstraNav-World: World Model for Foresight Control and Consistency
- 联合生成视觉与动作,双向约束提升预测可执行性
- 多步预测准确率提升,真实场景零样本成功率超90%
- 适合需要强泛化能力的机器人导航研究者
开放动态环境中具身导航需要对世界演化和动作序列进行精准预见。我们提出AstraNav-World,一种端到端世界模型,在统一的概率框架中联合推理未来视觉状态与动作序列。该框架融合基于扩散的视频生成器与视觉语言策略,实现预测场景与计划动作的同步演进。训练优化两个互补目标:生成条件于动作的多步视觉预测,以及基于预测视觉推导轨迹。这种双向约束使视觉预测可执行,确保决策扎根于物理一致、任务相关的未来,缓解解耦式“先设想后规划”流程中的累积误差问题。在多个具身导航基准测试中,轨迹精度与成功率达显著提升。消融实验表明紧密的视觉-动作耦合与联合训练不可或缺,任一分支移除均导致预测质量与策略可靠性下降。真实世界测试中,AstraNav-World展现出卓越零样本能力,无需任何现实微调即可适应未见过的场景。结果表明其捕捉到了可迁移的空间理解与规划相关导航动力学,而非仅过拟合仿真数据分布。总体而言,通过在单一生成模型中统一预见视觉与控制,我们朝着可靠、可解释且通用的具身智能体迈进,使其能在开放现实环境中稳健运行。
原文摘要 · Abstract (English)
Embodied navigation in open, dynamic environments demands accurate foresight of how the world will evolve and how actions will unfold over time. We propose AstraNav-World, an end-to-end world model that jointly reasons about future visual states and action sequences within a unified probabilistic framework. Our framework integrates a diffusion-based video generator with a vision-language policy, enabling synchronized rollouts where predicted scenes and planned actions are updated simultaneously. Training optimizes two complementary objectives: generating action-conditioned multi-step visual predictions and deriving trajectories conditioned on those predicted visuals. This bidirectional constraint makes visual predictions executable and keeps decisions grounded in physically consistent, task-relevant futures, mitigating cumulative errors common in decoupled "envision-then-plan" pipelines. Experiments across diverse embodied navigation benchmarks show improved trajectory accuracy and higher success rates. Ablations confirm the necessity of tight vision-action coupling and unified training, with either branch removal degrading both prediction quality and policy reliability. In real-world testing, AstraNav-World demonstrated exceptional zero-shot capabilities, adapting to previously unseen scenarios without any real-world fine-tuning. These results suggest that AstraNav-World captures transferable spatial understanding and planning-relevant navigation dynamics, rather than merely overfitting to simulation-specific data distribution. Overall, by unifying foresight vision and control within a single generative model, we move closer to reliable, interpretable, and general-purpose embodied agents that operate robustly in open-ended real-world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。