通过追踪表征到行为,揭示视觉-语言-动作模型的控制机制。
VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing

- 用跨模态和检查点漂移的核相关分析追踪表征演化。
- 发现不同模型在模态适配和动作解码上依赖不同路径。
- 适合研究模型可解释性与具身智能的开发者参考。
理解视觉-语言-动作(VLA)模型如何将多模态知识转化为具身控制仍是一个开放挑战。本文提出VLA-Trace,一个渐进式诊断框架,通过从表征动态到因果控制归因再到行为表现的统一证据链分析VLA模型。该框架结合跨模态与检查点漂移中心的核相关分析(CKA)追踪表征演化,采用注意力击穿干预识别模态特异性控制路径,并利用回溯级行为探针检验语义关联、捷径依赖与语义跟随能力。在π₀.₅和OpenVLA上的实验揭示三大发现:第一,两模型在VLA微调过程中表现出不同的模态特异性适应动态;第二,它们在动作解码阶段依赖不同的多模态路由策略与层间依赖关系;第三,尽管VLA策略在视觉引导轨迹生成上表现优异,但在细粒度语义跟随方面仍受限。这些发现指明了未来在表征保持适应、因果VLA电路与组合语义控制方面的研究方向。
原文摘要 · Abstract (English)
Understanding how Vision-Language-Action (VLA) models transform multimodal knowledge into embodied control remains an open challenge. We present VLA-Trace, a progressive diagnostic framework that analyzes VLA models through a unified evidence chain from representation dynamics to causal control attribution and behavioral manifestation. It specifically combines cross-modal and checkpoint-drift centered kernel alignment (CKA) to trace representation evolution, attention knockout interventions to identify modality-specific control pathways, and rollout-level behavioral probes to examine grounding, shortcut dependence, and semantic following. Experiments on $π_{0.5}$ and OpenVLA reveal three key findings. First, the two models exhibit distinct modality-specific adaptation dynamics during VLA finetuning. Second, they rely on different multimodal routing strategies and layer-wise dependencies during action decoding. Third, although VLA policies excel at visually grounded trajectory generation, they remain limited in fine-grained semantic following. These findings highlight future directions for representation-preserving adaptation, causal VLA circuits, and compositional semantic control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。