arXiv:2608.03563cs.RO2026-08中稿 · IROS 2026

用统一视觉运动目标提升视觉语言机器人模型的训练效率与鲁棒性

Unified Visuomotor Targets: Supervising VLAs Beyond Physical Actions

论文配图:Unified Visuomotor Targets: Supervising VLAs Beyond Physical Actions
图 1 · 摘自论文原文
  • 提出统一潜变量目标,联合预测运动控制与场景变化信息
  • 在仿真和真实双臂操作中显著提升训练效率与任务成功率
  • 无需修改结构或额外数据,特别适合资源受限场景

视觉语言动作(VLA)模型通常从视觉和语言观测中预测机器人动作。这一设计存在本质不匹配:视觉语言模型编码丰富的高层场景与目标表征,而机器人动作是低层次信号且任务结构有限。我们提出一种新思路:不改变架构,而是改变策略所预测的目标。UVT(统一视觉运动目标)是一种联合编码运动控制与视觉场景转移信息的统一潜变量预测目标,无需额外数据或架构调整。在两个代表性VLA系统上测试,涵盖仿真基准与真实双臂操作任务,UVT显著提升了训练效率、最终任务性能和策略鲁棒性,尤其在训练预算有限和环境复杂时表现突出。视频演示与更多定性结果见项目网页:https://unified-visuomotor-targets.github.io/

原文摘要 · Abstract (English)

VLA models are trained to predict robot actions from visual and language observations. This is a natural choice, but it creates a mismatch: VLMs encode rich, high-level representations of scenes and goals, while robot actions are low-level signals with limited task structure. We ask whether changing what the policy is trained to predict, rather than how it is architecturally designed, can yield better and more efficiently trained policies. We propose UVT (Unified Visuomotor Target), a unified latent prediction target that jointly encodes motor control and visual scene transition information, requiring no architectural changes and no additional data. Applied to two representative VLA systems across simulation benchmarks and real bimanual manipulation tasks, UVT improves training efficiency, final task performance, and policy robustness, with particularly strong gains under limited training budgets and challenging environmental conditions. Rollout videos and additional qualitative results are available at our project webpage: https://unified-visuomotor-targets.github.io/

视觉语言模型机器人控制强化学习训练效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。