arXiv:2607.05390cs.ROcs.CV2026-07中稿 · ECCV被引 1

构建大规模多视角视觉触觉数据集,助力可变形物体建模研究

Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models

论文配图:Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
图 1 · 摘自论文原文
  • 采用无标记视觉触觉三维追踪技术,捕捉物体全局运动与局部形变
  • 包含198种日常物品、1980组交互序列,超215小时多角度观测数据
  • 首次系统对比2D视频模型与3D粒子模型在可变形物体上的表现

预测物体动态(即世界建模)是机器人操作中的核心挑战,而可变形物体因其高维状态空间和复杂材料特性尤为困难。现有方法主要基于二维像素空间或显式三维几何空间学习动力学,但因缺乏多样化的大型真实世界数据,其优劣对比尚不清晰。为此,我们提出Deform360,一个大规模多视角视觉触觉数据集,涵盖198种日常物品、1,980组交互序列,以及来自41个环绕摄像头和双臂触觉夹具的超过215小时观测数据,用于捕捉全局运动与接触引发的局部形变。通过新颖的无标记视觉触觉三维追踪流程提取密集几何与运动信息,我们系统评估了当前最先进的世界模型,对比了2D视频模型与3D粒子模型的表现。最后,通过在可变形物体上执行机器人规划任务,初步验证了数据集的实际应用价值。分析揭示了结构先验与可扩展性之间的权衡,为未来通用可变形物体中心的世界建模研究提供了坚实基准。

原文摘要 · Abstract (English)

Predicting object dynamics (i.e., world modeling) is a fundamental challenge for robotic manipulation, and modeling deformable objects presents a particularly difficult case due to their high-dimensional state spaces and complex material properties. While current world models approach this through two distinct paradigms: learning the dynamics over the 2D pixel space or more explicit 3D geometric space. A systematic understanding of their relative strengths and limitations remains elusive due to the lack of diverse, large-scale real-world data. To address this, we present Deform360, a large-scale visuotactile dataset featuring 198 daily-life objects, 1,980 interaction sequences, and over 215 hours of observations from 41 surround-view cameras and bimanual tactile grippers to capture both global motion and contact-induced local deformations. Leveraging a novel markerless visuotactile 3D tracking pipeline to extract dense geometry and motion, we systematically evaluate current state-of-the-art world models, comparing 2D video models against 3D particle models. Finally, we provide a preliminary demonstration indicating the real-world applicability of our dataset by performing robot planning tasks on deformable objects. Our analysis reveals key insights into the trade-offs between structural priors and scalability, providing a solid benchmark for future research in generalizable deformable object-centric world modeling. Project website: https://deform360.lhy.xyz

可变形建模多模态数据机器人感知视觉触觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。