让机器人在执行任务时动态调整视觉重要性,提升动作预测准确率。
A Causality-aware Infer-diagnose-refine Framework for Test-time Modality Adaptation in VLA Models

- 通过因果推断分析视觉信息在不同阶段的重要性
- 测试时无需训练,显著提升多种模型的性能表现
- 适合需要实时适应环境变化的机器人操作场景
视觉-语言-动作(VLA)模型根据语言指令和视觉观察、本体感知状态预测序列动作。然而,如何融合多模态信息仍是开放问题,因为机器人操作包含长距离移动和近距离交互等动态阶段,视觉信息的重要性随时间变化。本文提出一种模型无关的推理-诊断-修正(IDR)框架,在测试时对VLA模型的动作预测进行优化。IDR首先在真实与反事实的视觉条件下推断动作,再通过因果效应估计动态重要性,最终以无训练方式修正动作。我们设计了基于零填充干预的反事实动作推断、基于范数的因果效应量化,以及门控残差融合机制实现该框架。在多个仿真基准和真实任务上的实验表明,该方法在多种VLA骨干模型上均取得性能提升,验证了测试时动态调整视觉重要性的有效性。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models predict sequential actions to execute tasks specified by language instructions, conditioned on visual observations and proprioceptive states. However, how to fuse modalities in VLA models remains an open problem, since robot manipulation involves dynamic phases, such as long-distance movements and close-range interactions, in which the importance of visual observations may vary over time. In this paper, we propose an infer-diagnose-refine (IDR) framework, a model-agnostic framework that can be integrated with diverse VLA architectures for refining action predictions at test time. IDR first infers actions under factual and counterfactual scenarios of visual observations, and then diagnoses the causal effects of visual observations as the estimated dynamic importance, which is finally used to refine the action predictions in a training-free manner. We further design a causality-aware action refiner to realize the IDR framework, including zero-padding interventions for inferring counterfactual actions, norm-based quantification for diagnosing causal effects, and gated residual fusion for refining actions. Extensive experiments on both simulation benchmarks and real-world tasks show improvements in overall performance across multiple VLA backbones, demonstrating the efficacy of dynamically adjusting visual importance at test time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。