arXiv:2605.00321cs.RO2026-05中稿 · ICML被引 2

通过干预方法检测视觉动作模型的因果错位,提升泛化可靠性。

Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models

论文配图:Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models
图 1 · 摘自论文原文
  • 引入干预性显著性分数(ISS),量化视觉区域对动作决策的因果影响。
  • 提出干扰质量比(NMR),数值越高表示任务无关特征影响越强。
  • 实验表明NMR可预测模型泛化表现,适用于机器人控制等具身智能场景。

视觉-语言-动作(VLA)策略在分布外情形下常失效,暗示其决策可能依赖于表面视觉相关性而非任务真实原因。本文将视觉-动作归因建模为干预估计问题,提出干预显著性得分(ISS)——一种用于估计视觉区域对动作预测因果影响的干预掩码方法,以及干扰质量比(NMR)——一个衡量归因于无关特征的标量指标。我们分析了ISS的统计性质,证明其可实现无偏估计,并刻画了动作预测误差作为因果影响代理的有效条件。在多种操作任务上的实验表明,NMR能有效预测模型泛化行为,且ISS生成的解释比现有方法更可信。结果表明,干预性归因提供了一种简单有效的诊断手段,用于识别具身策略中的因果错位。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) policies often fail under distribution shift, suggesting that decisions may depend on spurious visual correlations rather than task-relevant causes. We formulate visual-action attribution as an interventional estimation problem. Accordingly, we introduce the Interventional Significance Score (ISS), an interventional masking procedure for estimating the causal influence of visual regions on action predictions, and the Nuisance Mass Ratio (NMR), a scalar measure of attribution to task-irrelevant features. We analyze the statistical properties of ISS and show that it admits unbiased estimation, and we characterize conditions under which action prediction error provides a valid proxy for causal influence. Experiments across diverse manipulation tasks indicate that NMR predicts generalization behavior and that ISS yields more faithful explanations than existing interpretability methods. These results suggest that interventional attribution provides a simple diagnostic approach for identifying causal misalignment in embodied policies.

具身智能因果推理可解释性机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。