arXiv:2512.01952cs.CVcs.AI2025-12被引 4

让视频世界模型具备空间稳定性,提升导航可靠性。

GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment

  • 用几何与感知奖励自监督对齐预训练模型,增强空间一致性。
  • 在户外环境中导航轨迹更稳定,空间一致性优于监督微调。
  • 适合需要精准空间推理的机器人导航与虚拟环境应用。

近期视频世界建模进展使得大规模生成模型能够以高视觉保真度模拟具身环境,为预测、规划和控制提供强先验。然而,尽管模型逼真,其常缺乏几何定位,限制了对空间一致性要求高的导航任务使用。我们提出强化学习世界对齐(RLWG),一种自监督后训练框架,通过几何与感知奖励将预训练世界模型与可物理验证的结构对齐。类似于语言模型中的可验证反馈强化学习(RLVR),RLWG可利用多个衡量姿态循环一致性、深度重投影和时间一致性的奖励。我们以GrndCtrl为例,基于组相对策略优化(GRPO)实现奖励对齐,使模型保持稳定轨迹、一致几何结构和可靠推演,适用于具身导航。如同大语言模型的后训练对齐,GrndCtrl利用可验证奖励弥合生成预训练与具身行为之间的鸿沟,在室外环境中实现了优于监督微调的空间一致性和导航稳定性。

原文摘要 · Abstract (English)

Recent advances in video world modeling have enabled large-scale generative models to simulate embodied environments with high visual fidelity, providing strong priors for prediction, planning, and control. Yet, despite their realism, these models often lack geometric grounding, limiting their use in navigation tasks that require spatial coherence and stability. We introduce Reinforcement Learning with World Grounding (RLWG), a self-supervised post-training framework that aligns pretrained world models with a physically verifiable structure through geometric and perceptual rewards. Analogous to reinforcement learning from verifiable feedback (RLVR) in language models, RLWG can use multiple rewards that measure pose cycle-consistency, depth reprojection, and temporal coherence. We instantiate this framework with GrndCtrl, a reward-aligned adaptation method based on Group Relative Policy Optimization (GRPO), yielding world models that maintain stable trajectories, consistent geometry, and reliable rollouts for embodied navigation. Like post-training alignment in large language models, GrndCtrl leverages verifiable rewards to bridge generative pretraining and grounded behavior, achieving superior spatial coherence and navigation stability over supervised fine-tuning in outdoor environments.

世界模型强化学习导航几何对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。