解决视觉语言动作模型因视觉干扰导致指令失效的问题
CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention

- 通过反事实干预构建双路径去混淆图,分离语言与视觉干扰
- 在仿真和真实机器人上均提升52.3%成功率,显著改善泛化能力
- 适合关注机器人决策可靠性与因果推理的研究者
视觉-语言-动作(VLA)模型在机器人操作中取得显著进展,但普遍存在视觉覆盖现象。由于视觉流密集而语言指令稀疏,模型常因模态失衡陷入因果混淆,过度依赖突出物体或熟悉布局等视觉伪相关因素,忽略原始指令。为此,本文提出CofactVLA,一种基于反事实干预的因果去混淆框架。通过单次前向传播动态构建语言掩码的反事实分支,利用两种协同机制消除视觉干扰:一是动作级正交投影引导(OPG),在连续流匹配中将实际速度场从反事实视觉偏差中几何分离,提取纯语义意图;二是特征级反事实协方差缩减(CCR),通过惩罚协方差差异的正特征空间,显式抑制主导视觉捷径,保留因果语言意图。大量实验证明,CofactVLA在多个仿真基准上达到新SOTA。在真实机器人实验中,该方法有效缩小泛化差距,在分布外场景下实现52.3%的绝对成功率提升。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have driven significant progress in robotic manipulation, yet they fundamentally struggle with the vision-override phenomenon. Driven by the severe modality imbalance between dense visual streams and sparse linguistic instructions, VLAs frequently fall prey to causal confusion. Instead of treating language as the primary causal driver, the policy entirely bypasses the original instruction by overfitting to spurious visual confounders, such as prominent objects or familiar layouts. To systematically alleviate this bias, we formalize the process of action generation as a Dual-path Deconfounding Graph (DDG) and propose CofactVLA, a novel causal intervention framework. By dynamically constructing a language-masked counterfactual branch within a single forward pass, CofactVLA isolates and neutralizes visual confounders through two synergistic mechanisms. First, Action-Level Orthogonal Projection Guidance (OPG) geometrically projects the factual velocity field away from the counterfactual visual bias during continuous flow matching, extracting the pure semantic intent. Second, Feature-Level Counterfactual Covariance Reduction (CCR) mathematically deconfounds latent representations by penalizing the positive eigenspace of the covariance difference, explicitly suppressing dominant visual shortcuts while preserving the causal language intent. Extensive experiments demonstrate that CofactVLA establishes a new state-of-the-art across diverse simulation benchmarks. Beyond simulation, real-world robot experiments demonstrate the causal efficacy of our method in bridging the generalization gap, yielding a 52.3\% absolute success rate gain under out-of-distribution scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。