提出ReViP框架,解决视觉-语言-动作模型误判完成的问题。
ReViP: Mitigating False Completion in Vision-Language-Action Models with Vision-Proprioception Rebalance
- 引入视觉-本体觉再平衡机制,用进度感知视觉线索调节状态依赖。
- 在8个任务上实现26%成功率提升,显著减少虚假完成错误。
- 适合关注机器人操作鲁棒性与多模态融合的开发者和研究者。
视觉-语言-动作(VLA)模型通过融合视觉、语言和本体觉信息预测机器人动作,但现有方法直接融合本体信号,导致状态主导偏差和虚假完成问题。本文系统分析该问题,归因于模态失衡——策略过度依赖内部状态推进而忽视视觉证据。为此,我们构建首个虚假完成评估基准套件,包含8个任务及三种可控扰动(物体掉落、干扰物替换、重新布局),全面评估模型表现。同时提出新框架ReViP,通过外部任务阶段观察器提取进度感知视觉线索,动态调节语义感知与本体动态间的耦合关系。该机制增强环境感知,缓解状态驱动误差。大量实验表明,ReViP有效降低虚假完成率,在自建基准上较π₀模型提升26%,并在LIBERO、RoboTwin 2.0及真实场景中均取得显著改进。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have advanced robotic manipulation by combining vision, language, and proprioception to predict actions. However, previous methods fuse proprioceptive signals directly with vision-language features, resulting in state-dominant bias and \textbf{false completions} despite visible execution failures. We systematically analyze this failure mode, attributing it to modality imbalance, where policies overly rely on internal state progression and underuse visual evidence. To address this, we introduce the first \textbf{False-Completion Benchmark Suite}, featuring eight tasks with three controlled perturbations (\emph{Object Drop}, \emph{Distractor Swap}, \emph{Relayout}) to comprehensively evaluate false completion. Moreover, we propose \textbf{ReViP}, a novel VLA framework with \textbf{Vi}sion-\textbf{P}roprioception \textbf{Re}balance to enhance visual grounding and robustness under perturbations. The key insight is to introduce auxiliary \emph{progress-aware visual cues} to adaptively modulate the coupling between semantic perception and proprioceptive dynamics. Specifically, progress-aware visual cues are extracted by an external Task-Stage Observer, which performs task-relevant reasoning on real-time observations to drive task-stage feature-wise linear modulation, enhancing environmental awareness and mitigating state-driven errors. Extensive experiments show that ReViP effectively mitigates false completion and improves success rates over strong VLA baselines, achieving a \textbf{26\%} gain over $π_0$ model on our suite, with gains extending to LIBERO, RoboTwin 2.0, and real-world evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。