arXiv:2510.01711cs.ROcs.LG2025-10被引 15

让视觉语言动作模型更懂机器人状态,提升操控成功率

Contrastive Representation Regularization for Vision-Language-Action Models

  • 用机器人本体感知状态的相对距离做软监督,增强控制相关表征
  • 在真实机器人任务中将成功率从45.0%提升至58.3%
  • 无需改动训练流程,适配现有VLA模型,轻量高效

视觉-语言-动作(VLA)模型通过利用预训练视觉-语言模型(VLM)的丰富表征,在机器人操作任务中展现出强大能力。然而,其表征仍不理想,对控制动作和本体感知信息不够敏感。为此,我们提出机器人状态感知对比损失(RS-CL),一种简单有效的表征正则化方法,旨在弥合VLM表征与机器人信号之间的差距。具体而言,RS-CL利用状态间的相对距离作为软监督信号,使表征更贴近机器人的本体感知状态。该方法补充原有动作预测目标,增强了控制相关的表征学习,且轻量、完全兼容标准VLA训练流程。实验表明,RS-CL显著提升现有先进VLA模型性能:在RoboCasa-Kitchen基准上达到69.7%的最高水平,并在具有挑战性的真实机器人操作任务中将成功率从45.0%提升至58.3%。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown strong capabilities in robot manipulation by leveraging rich representations from pre-trained Vision-Language Models (VLMs). However, their representations arguably remain suboptimal, lacking sensitivity to robotic signals such as control actions and proprioceptive information. To address the issue, we introduce Robot State-aware Contrastive Loss (RS-CL), a simple and effective representation regularization for VLA models, designed to bridge the gap between VLM representations and robotic signals. In particular, RS-CL aligns the representations more closely with the robot's proprioceptive states by using relative distances between the states as soft supervision. Complementing the original action prediction objective, RS-CL enhances control-relevant representation learning, while being lightweight and fully compatible with standard VLA training pipelines. Our empirical results demonstrate that RS-CL substantially improves the performance of state-of-the-art VLA models; it pushes the prior art to 69.7% achieving the state-of-the-art performance on the RoboCasa-Kitchen benchmark, and boosts success rates from 45.0% to 58.3% on challenging real-robot manipulation tasks.

机器人操控表征学习对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。