arXiv:2608.29967cs.ROcs.AI2026-08

用语言反馈修正视觉语言动作模型执行偏差,无需重训练。

Training-Free Action Correction for VLA Model Failures via Language Feedback

论文配图:Training-Free Action Correction for VLA Model Failures via Language Feedback
图 1 · 摘自论文原文
  • 通过自然语言指令调整动作幅度,不改变模型权重。
  • 实机实验中使机器人成功率从几乎为零恢复至接近完美。
  • 仅对策略正确但动作不准的错误有效,适合部署时快速修复。

视觉-语言-动作(VLA)模型具备强大的语义理解能力,但在部署中仍存在系统性失败。这些失败发生条件及是否可无需重训练修复尚不明确。本文提出CorrectVLA框架,将任务级自然语言修正转化为不修改策略权重的附加动作幅度调整。人类提供单一任务级修正,统一应用于所有推演过程,无需每轮干预。在仿真中,CorrectVLA成功修复分布内与分布外任务的执行错位问题。在真实机器人(UFactory xArm7)上,面对环境变化,基础策略几乎完全失效时,CorrectVLA使其成功率恢复至接近完美,且泛化至不同物体位置与身份。基于LIBERO-90的失败模式分类表明,仅当策略能正确理解目标但动作幅度校准错误时,该方法才有效;若语义理解本身失效,则无法修复。该方法在策略具备战略正确性时成功,而根本理解缺失时失败,确立了推理阶段修正的实用边界。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models demonstrate strong semantic understanding yet exhibit systematic failures during deployment. The conditions under which these failures occur, and whether they can be corrected without retraining, remain poorly understood. In this paper, we take steps toward addressing this gap. We present CorrectVLA, a framework that translates task-level natural language corrections into additive action magnitude adjustments without modifying policy weights. A human provides a single task-level correction, applied uniformly across all rollouts without per-episode intervention. In simulation, CorrectVLA recovers execution misalignment failures across both in-distribution and OOD tasks. In real-robot experiments on a UFactory xArm7 under environment shift, CorrectVLA restores near-perfect success where the base policy almost entirely breaks down, generalizing across object locations and identities. Through a taxonomy of failure modes on LIBERO-90, we find that execution misalignment failures, where the policy reaches the correct target but miscalibrates action magnitudes, represent the correctable subset, while other failure modes where semantic comprehension itself breaks down are not amenable to this approach. The approach succeeds when policies possess strategic correctness and fails when fundamental comprehension is absent, establishing a practical operational boundary for inference-time correction.

视觉语言动作机器人在线修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。