让视觉语言动作模型精准抓取并自动判断任务完成,减少失败和超时。
From Knowing to Doing Precisely: A General Self-Correction and Termination Framework for VLA models
- 引入自校正控制循环,动态修正动作偏差。
- 在LIBERO基准上提升所有任务成功率,减少冗余动作。
- 无需训练、轻量级,适合部署于复杂环境的机器人系统。
尽管用于具身智能体的视觉-语言-动作(VLA)模型整合了感知、推理与控制,仍存在两大关键缺陷:其一,在抓取任务中,语言模型生成的动作标记常出现细微空间偏差,导致抓取失败;其二,缺乏可靠的任务完成识别能力,引发冗余动作与频繁超时错误。为克服这些挑战并提升鲁棒性,我们提出一种轻量、免训练的框架VLA-SCT。该框架作为自校正控制回路,结合数据驱动的动作精修与条件逻辑终止机制。相比基线方法,本方法在LIBERO基准的所有数据集上均实现稳定提升,显著提高精细操作任务的成功率,并确保准确的任务完成判定,从而推动更可靠的VLA智能体在复杂、非结构化环境中的应用。
原文摘要 · Abstract (English)
While vision-language-action (VLA) models for embodied agents integrate perception, reasoning, and control, they remain constrained by two critical weaknesses: first, during grasping tasks, the action tokens generated by the language model often exhibit subtle spatial deviations from the target object, resulting in grasp failures; second, they lack the ability to reliably recognize task completion, which leads to redundant actions and frequent timeout errors. To address these challenges and enhance robustness, we propose a lightweight, training-free framework, VLA-SCT. This framework operates as a self-correcting control loop, combining data-driven action refinement with conditional logic for termination. Consequently, compared to baseline approaches, our method achieves consistent improvements across all datasets in the LIBERO benchmark, significantly increasing the success rate of fine manipulation tasks and ensuring accurate task completion, thereby promoting the deployment of more reliable VLA agents in complex, unstructured environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。