arXiv:2509.07962cs.RO2025-09中稿 · CoRL被引 32

让机器人模型学会感知扭矩,提升物理交互能力

TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models

  • 在解码器中加扭矩适配器,效果优于编码器
  • 预测扭矩作为辅助输出,任务成功率提升显著
  • 适合研究机器人具身智能与物理感知的学者

许多机器人操作任务需要感知和响应力信号(如扭矩),以判断任务是否成功并实现闭环控制。然而,现有视觉-语言-动作(VLA)模型缺乏整合此类细微物理反馈的能力。本文系统探索了扭矩感知型VLA模型的设计空间,提出多种集成扭矩信号的策略。实验发现:在解码器中引入扭矩适配器,性能显著优于在编码器中插入;此外,受自动驾驶中联合预测与规划范式启发,提出将扭矩预测作为辅助输出,促使模型建立对交互动力学的物理基础表征,进一步提升性能。在多个接触丰富的操作基准上进行了充分的定量与定性实验,验证了方法的有效性。

原文摘要 · Abstract (English)

Many robotic manipulation tasks require sensing and responding to force signals such as torque to assess whether the task has been successfully completed and to enable closed-loop control. However, current Vision-Language-Action (VLA) models lack the ability to integrate such subtle physical feedback. In this work, we explore Torque-aware VLA models, aiming to bridge this gap by systematically studying the design space for incorporating torque signals into existing VLA architectures. We identify and evaluate several strategies, leading to three key findings. First, introducing torque adapters into the decoder consistently outperforms inserting them into the encoder.Third, inspired by joint prediction and planning paradigms in autonomous driving, we propose predicting torque as an auxiliary output, which further improves performance. This strategy encourages the model to build a physically grounded internal representation of interaction dynamics. Extensive quantitative and qualitative experiments across contact-rich manipulation benchmarks validate our findings.

机器人物理感知多模态扭矩建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。