无需训练即可实时修复视觉语言动作模型的抓取失败
ProbeAct: Probe-Guided Training-Free Failure Recovery in Vision-Language-Action Models

- 用轻量级探针从中间特征预测物体3D位置,支持多物体跟踪
- 仅靠夹爪内部信号和末端执行器运动检测抓取、搬运、放置失败
- 通过安全约束动态修正动作,不改变原模型且适配各类预训练模型
视觉-语言-动作(VLA)模型在训练分布内表现出色,但面对光照变化、视角调整或初始状态微小差异时易失效。本文提出PROBEACT,一种无需训练的运行时干预框架,可在不修改权重或额外示范的前提下,检测并恢复预训练VLA策略中的抓取与放置失败。该框架包含三部分:(i) 轻量级多目标隐状态探针,从VLA中间特征预测任务相关物体3D位置,并通过匈牙利匹配实现多物体身份追踪;(ii) 无对象依赖的运动学状态机,仅利用夹爪内部信号和末端执行器运动检测抓取、运输、放置失败;(iii) 分层控制屏障函数(CBF)滤波器,将重复失败区域编码为软安全集约束,最小化修正动作同时保持原始行为。在LIBERO-plus基准测试中,该框架使OpenVLA-OFT模型成功率从69.6%提升至74.1%,对基础和微调后的VLA策略均具广泛适用性。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models demonstrate strong perfor-1 mance on language-conditioned robotic manipulation within their training dis-2 tribution, yet their generalization capabilities remain fundamentally limited. They3 lack the robustness required to handle perturbations, frequently failing when con-4 fronted with lighting changes, altered camera viewpoints, or small initial-state5 variations. We propose PROBEACT, a training-free runtime intervention frame-6 work that detects and recovers from grasping and placement failures in pre-7 trained VLA policies without modifying their weights or requiring additional8 demonstrations. PROBEACT combines three components: (i) a lightweight multi-9 target hidden-state probe that predicts the 3D positions of task-relevant objects10 from intermediate VLA features, with Hungarian-matched identity tracking for11 multi-object scenes; (ii) an object-agnostic kinematic state machine that detects12 grasp, transport, and placement failures using only gripper-internal signals and13 end-effector kinematics; and (iii) a hierarchical Control Barrier Function (CBF)14 filter that encodes repeated-failure locations as soft safe-set constraints, mini-15 mally correcting VLA actions while preserving baseline behavior. As a plug-and-16 play, training-free intervention loop, PROBEACT is orthogonal to existing train-17 ing pipelines. Evaluated on the LIBERO-plus benchmark, our framework acts as18 a universal safety net, improving the success rate of the OpenVLA-OFT model19 from 69.6% to 74.1%, while demonstrating broad applicability to both base and20 fine-tuned VLA policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。