arXiv:2606.04825cs.RO2026-06

HapTile构建了融合触觉反馈的多模态操作数据集,提升机器人抓取稳定性。

HapTile: A Haptic-Informed Vision-Tactile-Language-Action Dataset for Contact-Rich Imitation Learning

论文配图:HapTile: A Haptic-Informed Vision-Tactile-Language-Action Dataset for Contact-Rich Imitation Learning
图 1 · 摘自论文原文
  • 在遥操作中加入指尖触觉反馈,让操作员实时感知接触力
  • 涵盖10类日常任务,包含语言指令、视觉、触觉与动作轨迹同步数据
  • 适合研究触觉增强型模仿学习的科研人员使用

尽管触觉感知对可靠操作至关重要,现有视觉-语言-动作(VLA)数据集大多仅依赖视觉信息,而少数含触觉的数据集缺乏任务多样性、语言条件和动作轨迹的联合。此外,现有远程操控系统很少提供触觉反馈,但已有研究表明触觉能显著提升示范质量和操作稳定性。本文提出HapTile,一个基于物理交互的多模态操作数据集,通过两种层级的触觉感知实现突破:机器人末端执行器的指尖触觉传感器采集接触信息,以及远程操作端的触觉反馈机制。数据采集平台将触觉反馈直接嵌入遥控控制器,使操作者可实时感知接触力。系统基于标准可复现的机器人平台,并配备自研指尖触觉传感器。数据集包含拾取放置、折叠、按压、堆叠等10类高接触密度任务,每项任务配有语言指令,同步记录视觉、触觉观测与动作轨迹。我们还提供了两个基线模型的基准测试,验证该数据集在高接触任务策略学习中的有效性。数据集及详细信息见 haptile-dataset.github.io。

原文摘要 · Abstract (English)

Despite the importance of tactile sensing for reliable manipulation, most existing Vision-Language-Action (VLA) datasets remain vision-only, and those that do incorporate tactile information typically lack the joint combination of task diversity, language conditioning, and action trajectories. Furthermore, existing teleoperation pipelines rarely provide haptic feedback to the operator, despite its established role in demonstration quality and manipulation stability. In this work, we present HapTile, a contact-grounded visuotactile manipulation dataset that advances beyond vision-only trajectory datasets by embedding physical interaction sensing at two levels: fingertip tactile feedback at the robot end-effector, and haptic-informed demonstrations at the teleoperator side. The data collection platform integrates haptic feedback directly into the teleoperation controller, enabling the operator to perceive contact interactions in real time. It is built around a standard and reproducible robotic system equipped with custom-designed fingertip tactile sensors. The dataset comprises everyday manipulation tasks spanning a broad range of contact-rich skills, including pick-and-place, folding, pressing, stacking, and other routine activities. Each task is paired with language instructions that condition the policy on the manipulation objective, together with synchronized visuotactile observations and action trajectories. In addition, we provide a benchmarking study on contact-rich policy learning using two baseline models to evaluate the effectiveness of the proposed contact-grounded dataset. The dataset and additional details are available on our website: haptile-dataset.github.io.

触觉感知模仿学习多模态数据集机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。