arXiv:2603.15257cs.RO2026-03被引 12

让机器人无需触觉传感器也能精准完成接触密集操作

HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing

  • 用预训练的触觉奖励指导动作生成,离线学习触觉感知能力
  • 真实场景中成功率86.7%,优于带触觉反馈的基线模型
  • 适合希望降低成本、提升部署通用性的机器人研发者

触觉感知对视觉-语言-动作(VLA)系统在接触密集任务中的灵巧与安全操作至关重要。然而,依赖专用触觉硬件会增加成本并降低跨平台可复现性。本文认为,触觉感知能力可在离线阶段学习,并在推理时无需直接触觉反馈即可部署。为此提出HapticVLA,包含两个紧密耦合阶段:安全感知的奖励加权流匹配(SA-RWFM)与触觉蒸馏(TD)。SA-RWFM训练一个基于流匹配的动作专家,引入预计算的安全感知触觉奖励,惩罚过大的抓握力和次优抓握轨迹。TD将该触觉感知能力蒸馏至常规VLA:从SA-RWFM教师模型中提取紧凑触觉令牌,并训练学生VLA仅通过视觉与状态模态预测该令牌,从而在推理时实现触觉感知动作生成,无需机载触觉传感器。此设计在保留接触密集任务中触觉感知推理能力的同时,去除了部署阶段对触觉传感器的需求。真实世界实验显示,HapticVLA平均成功率达86.7%,持续优于基线VLA——包括在推理时提供直接触觉反馈的版本。

原文摘要 · Abstract (English)

Tactile sensing is a crucial capability for Vision-Language-Action (VLA) architectures, as it enables dexterous and safe manipulation in contact-rich tasks. However, reliance on dedicated tactile hardware increases cost and reduces reproducibility across robotic platforms. We argue that tactile-aware manipulation can be learned offline and deployed without direct haptic feedback at inference. To this end, we present HapticVLA, which proceeds in two tightly coupled stages: Safety-Aware Reward-Weighted Flow Matching (SA-RWFM) and Tactile Distillation (TD). SA-RWFM trains a flow-matching action expert that incorporates precomputed, safety-aware tactile rewards penalizing excessive grasping force and suboptimal grasping trajectories. TD further transfers this tactile-aware capability into a conventional VLA: we distill a compact tactile token from the SA-RWFM teacher and train a student VLA to predict that token from vision and state modalities, enabling tactile-aware action generation at inference without requiring on-board tactile sensors. This design preserves contact-rich tactile-aware reasoning within VLA while removing the need for on-board tactile sensors during deployment. On real-world experiments, HapticVLA achieves a mean success rate of 86.7%, consistently outperforming baseline VLAs - including versions provided with direct tactile feedback during inference.

触觉感知机器人操作视觉语言动作无传感器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。