通过视觉与参数空间协同,提升机器人长时序模拟的准确性与一致性。
ViPSim: Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Models

- 融合像素对齐的视觉先验与动作数值驱动,实现双空间协同建模。
- 在长时滚动中显著减少轨迹漂移,支持复杂柔体交互如折叠布料。
- 适用于跨任务、跨机器人形态的高保真评估,适合智能体训练与验证。
具身世界模型(EWMs)作为推进具身智能的可扩展且无风险范式,能安全评估视觉-语言-动作系统。然而,其可靠性常受低维动作与高维视频生成间表征差距制约,导致几何对应缺失,表现为长时滚动中轨迹漂移累积和机器人-物体交互不一致。为此,我们提出ViPSim框架,通过视觉空间与参数空间的协同合作实现一致的长时序生成。视觉空间引入显式的空间先验,整合末端执行器位姿的像素对齐投影、相机视角、深度感知场景几何及机器人形态掩码,提供密集结构约束;参数空间则注入原始动作序列与相机矩阵,提供精确运动引导。两者统一后,生成状态同时受几何边界锚定和数值指令驱动。大量实验表明,ViPSim具有模型无关性,显著提升轨迹一致性。其在柔性物体复杂交互(如布料折叠)中展现涌现能力,并在分布外与跨具身场景中保持鲁棒性能,为具身智能体的自动化评估与预测控制提供高保真基础。
原文摘要 · Abstract (English)
Embodied World Models (EWMs) have emerged as a scalable and risk-free paradigm for advancing embodied intelligence, enabling the safety-critical evaluation of Vision-Language-Action systems. However, their reliability as evaluation benchmarks and foundational simulators is often hindered by the representation gap between low-dimensional actions and high-dimensional video synthesis. This gap results in a lack of geometric correspondence, manifesting as accumulated trajectory drift and inconsistent robot-object interactions during long-horizon rollouts. To bridge this gap, we propose ViPSim, a framework that achieves consistent long-horizon generation through the synergistic collaboration of Visual and Parameter Spaces. We define the Visual Space as a domain of explicit spatial priors, integrating pixel-aligned projections of end-effector pose, camera perspectives, depth-informed scene geometry, and robotic morphological masks to provide dense structural grounding. Concurrently, the Parameter Space serves as a domain of numerical drivers, injecting raw action sequences and camera matrices to provide precise motion guidance. By unifying these two spaces, ViPSim ensures that the generated states are simultaneously anchored by geometric boundaries and steered by numerical commands. Extensive experiments demonstrate that ViPSim is backbone-agnostic and significantly enhances trajectory consistency. Notably, our approach exhibits emergent capabilities in generating complex interactions with deformable objects (e.g., cloth folding) and maintains robust performance in out-of-distribution and cross-embodiment scenarios, providing a high-fidelity foundation for the automated evaluation and predictive control of embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。