arXiv:2609.03889cs.ROcs.AI2026-09

让机器人在复杂操作中感知并补偿接触力,无需额外传感器。

FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation

论文配图:FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation
图 1 · 摘自论文原文
  • 用无传感器方法估计接触力及其变化,注入视觉语言动作模型
  • 在5000+真实任务数据上微调模型,实现精准力感知与动作生成
  • 适合需要高精度物理交互的轮腿机器人应用场景

接触丰富的运动-操作任务需要连接语义动作生成与物理交互控制。现有视觉-语言-动作(VLA)模型从视觉和语言输入生成任务级动作,但无法解释这些动作引发的物理交互。虽然全身控制(WBC)策略能稳定机器人,却难以区分任务相关的交互力与外部干扰力。尽管力/力矩传感器可直接测量物理交互,但改造需额外硬件成本和集成工作,尤其对未设计传感集成的平台。为此,我们提出FWBC-VLA,一种融合任务级VLA动作生成与底层全身补偿控制的力感知框架。首先引入HSR-Force,一种无传感器的残余力矩估计器,用于推断接触强度及其时间变化。这些接触估计以标记形式注入VLA动作专家的解码过程,使策略能够感知接触开始、持续受力和释放。针对运动-操作任务,所有预训练VLA主干参数在包含超过5000个任务实例的WL&Arm数据集上微调。同时,将机器人的本体感知状态、基于雅可比矩阵的机体坐标系力估计及估计的接触状态联合输入补偿生成器,生成修正动作。最终,操作中心动作与修正动作结合后送入WBC策略执行。白板清洁与带门缓闭器的开门等真实世界实验验证了该框架在接触丰富运动-操作中的有效性。

原文摘要 · Abstract (English)

Contact-rich loco-manipulation requires a bridge between semantic action generation and physical interaction control. Existing Vision-language-action (VLA) models generate task-level actions from visual and linguistic observations, but cannot interpret the physical interactions induced by those actions. While the whole-body control (WBC) policy can stabilize the robot, it cannot distinguish task-relevant interaction forces from forces induced by external disturbances during manipulation. Although force/torque sensors provide direct measurements of physical interactions, retrofitting them entails additional hardware costs and substantial integration effort, particularly for platforms not designed with sensor integration in mind. To address this problem, we propose FWBC-VLA, a force-aware framework that bridges task-level VLA action generation and low-level whole-body compensation control for wheeled-legged robots. First, we introduce HSR-Force, a sensorless residual-torque estimator for inferring contact strength and its temporal variation. These contact estimates are then encoded as tokens and injected into the VLA action expert during action decoding, enabling the policy to perceive contact onset, sustained loading, and release. For loco-manipulation tasks, all parameters of the pretrained VLA backbone are fine-tuned on our WL\&Arm Dataset, which comprises more than 5,000 episodes. Moreover, the robot's proprioceptive state, the Jacobian-derived body-frame force estimate, and the estimated contact state are jointly fed into a compensation generator to produce corrective actions. The manipulation-centric actions are subsequently combined with the corrective actions and passed to the WBC policy for execution. Real-world experiments on whiteboard wiping and door opening with a door closer demonstrate the effectiveness of our FWBC-VLA in contact-rich loco-manipulation.

力感知机器人控制轮腿机器人无传感器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。