arXiv:2603.20101cs.AI2026-03被引 2

用自动化代理分析模型组件作用,发现评估方法存在重大缺陷。

Pitfalls in Evaluating Interpretability Agents

  • 构建迭代实验的智能代理,自动分析模型电路功能。
  • 对比六项任务中专家解释,系统表现看似接近人类水平。
  • 揭示复制评估的三大陷阱,提出基于功能互换性的新评测法。

自动化可解释性系统旨在减少人工依赖,扩展至更大模型和多样任务。近期研究利用大语言模型(LLMs)实现从固定流程到完全自主解释代理的演进。这一转变要求评估方法同步提升以应对生成解释的数量与复杂性。本文聚焦自动化电路分析——解释模型组件在特定任务中的角色。我们构建了一个研究代理,可迭代设计实验并修正假设。在六项文献中的电路分析任务上,该系统与人类专家解释相比表现相当。然而深入分析发现,基于复现的评估存在多重缺陷:专家解释可能存在主观或不完整,结果导向比对掩盖了研究过程,而基于LLM的系统可能通过记忆或推测复现已有成果。为此,我们提出一种无监督的内在评估方法,基于模型组件的功能互换性。本工作揭示了复杂自动化可解释性系统评估的根本挑战,指出了复制评估的关键局限。

原文摘要 · Abstract (English)

Automated interpretability systems aim to reduce the need for human labor and scale analysis to increasingly large models and diverse tasks. Recent efforts toward this goal leverage large language models (LLMs) at increasing levels of autonomy, ranging from fixed one-shot workflows to fully autonomous interpretability agents. This shift creates a corresponding need to scale evaluation approaches to keep pace with both the volume and complexity of generated explanations. We investigate this challenge in the context of automated circuit analysis -- explaining the roles of model components when performing specific tasks. To this end, we build an agentic system in which a research agent iteratively designs experiments and refines hypotheses. When evaluated against human expert explanations across six circuit analysis tasks in the literature, the system appears competitive. However, closer examination reveals several pitfalls of replication-based evaluation: human expert explanations can be subjective or incomplete, outcome-based comparisons obscure the research process, and LLM-based systems may reproduce published findings via memorization or informed guessing. To address some of these pitfalls, we propose an unsupervised intrinsic evaluation based on the functional interchangeability of model components. Our work demonstrates fundamental challenges in evaluating complex automated interpretability systems and reveals key limitations of replication-based evaluation.

可解释性自动化评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。