提出视觉触觉世界模型基准,揭示结构化表征对复杂操作规划的关键作用。
ContactWorld: What Matters in Vision-Tactile World Models for Contact-Rich Manipulation

- 采用点云+触觉力场融合表征,提升空间与时间连续性
- 点云观测使规划成功率从20.7%提升至32.1%,结合触觉达36.1%
- 长期规划中触觉重要性凸显,需关注多模态兼容性
接触丰富操作需要世界模型从多模态感知中推理复杂的接触动力学。然而,何种表征特性能支持接触密集场景下的稳定长时程规划仍不明确。本文提出ContactWorld,一个涵盖12项接触密集操作任务(包括插入、拆解、拧紧和探索交互)的基准与系统性实证研究。大量实验表明,兼具空间结构与时间连续性的表征始终表现最优。具体而言,点云观测将平均规划成功率从腕部视角的20.7%、前视视角的22.0%提升至32.1%。进一步发现,触觉的有效性关键取决于跨模态表征兼容性,而非单纯模态扩展。结合点云与触觉力场表示(保留更丰富的空间结构与交互动态),性能进一步提升至36.1%,在所有任务中表现最佳。此外,在长时程规划目标下,触觉作用愈发显著,因预测误差与接触不确定性随时间累积。上述结果强调了表征结构、多模态兼容性与长时程鲁棒性在视觉触觉世界模型中的核心价值。
原文摘要 · Abstract (English)
Contact-rich manipulation requires world models to reason over complex contact dynamics from multimodal sensory observations. However, it remains unclear which representation properties fundamentally support stable long-horizon planning in contact-rich settings. In this paper, we present ContactWorld, a benchmark and systematic empirical study of vision-tactile world models spanning 12 contact-rich manipulation tasks, including insertion, disassembly, screwing, and exploratory interaction. Across extensive experiments, we find that representations that are both spatially structured and temporally continuous consistently achieve the strongest planning performance. In particular, point-cloud observations improve average planning success rates from 20.7% with wrist-view observations and 22.0% with front-view observations to 32.1%. We further find that the effectiveness of tactile sensing depends critically on cross-modal representation compatibility rather than modality scaling alone. Combining point-cloud observations with tactile force-field representations, which preserve richer spatial structure and interaction dynamics, further improves performance to 36.1%, yielding the strongest overall planning performance across all evaluated tasks. Moreover, tactile sensing becomes increasingly important under long-horizon planning objectives, where compounding prediction errors and contact uncertainty accumulate over time. Together, these findings highlight the importance of representation structure, multimodal compatibility, and long-horizon robustness in vision-tactile world models for contact-rich robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。