arXiv:2604.20444cs.ROcs.AI2026-04被引 1

构建多模态触觉数据集,提升双手操作任务的物理交互感知能力

VTouch++: A Multimodal Dataset with Vision-Based Tactile Enhancement for Bimanual Manipulation

论文配图:VTouch++: A Multimodal Dataset with Vision-Based Tactile Enhancement for Bimanual Manipulation
图 1 · 摘自论文原文
  • 基于视觉的触觉传感技术,捕捉高保真物理交互信号
  • 采用矩阵式任务设计,支持系统化学习与评估
  • 覆盖真实场景的自动化数据采集,适配多机器人通用推理

近年来,具身智能发展迅速,但双臂操作尤其是高接触密度的任务仍面临挑战。主要源于缺乏包含丰富物理交互信号、系统化任务结构和足够规模的数据集。为此,我们提出VTOUCH++数据集:利用视觉驱动的触觉感知技术提供高保真物理交互信号,采用矩阵式任务设计支持系统性学习,并通过自动化数据采集流程覆盖真实世界中需求驱动的场景,确保可扩展性。为验证数据集有效性,我们在跨模态检索和真实机器人评估上进行了广泛定量实验。最终,展示了在多个机器人、策略和任务间具备泛化能力的实际表现。

原文摘要 · Abstract (English)

Embodied intelligence has advanced rapidly in recent years; however, bimanual manipulation-especially in contact-rich tasks remains challenging. This is largely due to the lack of datasets with rich physical interaction signals, systematic task organization, and sufficient scale. To address these limitations, we introduce the VTOUCH dataset. It leverages vision based tactile sensing to provide high-fidelity physical interaction signals, adopts a matrix-style task design to enable systematic learning, and employs automated data collection pipelines covering real-world, demand-driven scenarios to ensure scalability. To further validate the effectiveness of the dataset, we conduct extensive quantitative experiments on cross-modal retrieval as well as real-robot evaluation. Finally, we demonstrate real-world performance through generalizable inference across multiple robots, policies, and tasks.

多模态触觉感知双臂操作数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。