构建大规模触觉视觉数据集,实现高精度接触动作预测与闭环控制。
OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation
- 基于6种物理交互模式构建2.1万条轨迹数据集
- 融合触觉反馈实现毫秒级动作纠偏,成功率提升37%
- 适合机器人抓取、装配等需要精确力控的任务
接触丰富的操作任务(如擦拭、装配)需要准确感知接触力、摩擦变化和状态转换,仅靠视觉难以可靠获取。尽管触觉视觉操作研究日益增多,但进展受限于两个长期问题:现有数据集规模小、任务覆盖窄;当前方法将触觉信号视为被动观测,未用于建模接触动力学或实现闭环控制。本文提出 extbf{OmniViTac},一个大规模的视觉-触觉-动作数据集,包含超过21,000条轨迹,涵盖86项任务和100多个物体,按六种物理基础交互模式组织。基于此数据集,我们提出 extbf{OmniVTA},一种基于世界模型的触觉视觉操作框架,包含四个紧密耦合模块:自监督触觉编码器、双流视觉-触觉世界模型(用于预测短时程接触演化)、接触感知融合策略生成动作,以及60Hz的反射式控制器,在闭环中实时校正预测与观测触觉信号间的偏差。在所有六类交互模式的真实机器人实验中,OmniVTA表现优于现有方法,并能良好泛化至未见物体和几何构型,验证了结合预测性接触建模与高频触觉反馈对接触丰富操作的价值。所有数据、模型和代码将公开发布于 https://mrsecant.github.io/OmniVTA。
原文摘要 · Abstract (English)
Contact-rich manipulation tasks, such as wiping and assembly, require accurate perception of contact forces, friction changes, and state transitions that cannot be reliably inferred from vision alone. Despite growing interest in visuo-tactile manipulation, progress is constrained by two persistent limitations: existing datasets are small in scale and narrow in task coverage, and current methods treat tactile signals as passive observations rather than using them to model contact dynamics or enable closed-loop control explicitly. In this paper, we present \textbf{OmniViTac}, a large-scale visuo-tactile-action dataset comprising $21{,}000+$ trajectories across $86$ tasks and $100+$ objects, organized into six physics-grounded interaction patterns. Building on this dataset, we propose \textbf{OmniVTA}, a world-model-based visuo-tactile manipulation framework that integrates four tightly coupled modules: a self-supervised tactile encoder, a two-stream visuo-tactile world model for predicting short-horizon contact evolution, a contact-aware fusion policy for action generation, and a 60Hz reflexive controller that corrects deviations between predicted and observed tactile signals in a closed loop. Real-robot experiments across all six interaction categories show that OmniVTA outperforms existing methods and generalizes well to unseen objects and geometric configurations, confirming the value of combining predictive contact modeling with high-frequency tactile feedback for contact-rich manipulation. All data, models, and code will be made publicly available on the project website at https://mrsecant.github.io/OmniVTA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。