视频模型内部存在处理动作成败的隐藏计算机制。
Attention Gathers, MLPs Compose: A Causal Analysis of an Action-Outcome Circuit in VideoViT
- 注意力头负责收集低层证据,MLP块负责生成成功信号
- 从第5到11层逐步放大成败信号,形成级联放大结构
- 适合关注模型可解释性与可信AI的研究者
本文通过可解释人工智能方法,特别是机制可解释性技术,逆向解析了一个预训练视频视觉变换器中表征动作结果的内部电路。研究发现,'成功与失败'信号通过一个独特的放大级联过程实现。尽管第0层已有低层差异,但抽象语义层面的结果表示在第5至11层逐步增强。因果分析(主要基于激活补丁和消融实验)揭示了明确分工:注意力头充当'证据收集者',提供部分信号恢复所需信息;而MLP块作为强健的'概念组合者',是生成'成功'信号的主要驱动因素。该分布式且冗余的内部电路解释了模型对简单消融的鲁棒性,展现出处理人类动作结果的核心计算模式。关键在于,即使仅训练用于简单分类任务的模型,也发展出超越显式任务的复杂结果表征能力,表明模型可能具备'隐含知识',凸显了构建真正可解释、可信AI系统所需的机制监督必要性。
原文摘要 · Abstract (English)
The paper explores how video models trained for classification tasks represent nuanced, hidden semantic information that may not affect the final outcome, a key challenge for Trustworthy AI models. Through Explainable and Interpretable AI methods, specifically mechanistic interpretability techniques, the internal circuit responsible for representing the action's outcome is reverse-engineered in a pre-trained video vision transformer, revealing that the "Success vs Failure" signal is computed through a distinct amplification cascade. While there are low-level differences observed from layer 0, the abstract and semantic representation of the outcome is progressively amplified from layers 5 through 11. Causal analysis, primarily using activation patching supported by ablation results, reveals a clear division of labor: Attention Heads act as "evidence gatherers", providing necessary low-level information for partial signal recovery, while MLP Blocks function as robust "concept composers", each of which is the primary driver to generate the "success" signal. This distributed and redundant circuit in the model's internals explains its resilience to simple ablations, demonstrating a core computational pattern for processing human-action outcomes. Crucially, the existence of this sophisticated circuit for representing complex outcomes, even within a model trained only for simple classification, highlights the potential for models to develop forms of 'hidden knowledge' beyond their explicit task, underscoring the need for mechanistic oversight for building genuinely Explainable and Trustworthy AI systems intended for deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。