arXiv:2511.18409cs.CLcs.AI2025-11被引 1

评测大模型可解释性技术,发现有效定位关键组件与变量的方法。

Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models

  • 通过集成与正则化提升电路定位准确率
  • 低维非线性投影显著改善特征映射效果
  • 面向社区开放评测框架,适合可解释性研究者

机制可解释性旨在揭示语言模型如何实现特定行为,但其进展评估仍具挑战。新发布的机制可解释性基准(MIB;Mueller等,2025)提供了标准化评估框架。基于此,BlackboxNLP 2025 共享任务将 MIB 扩展为全社区可复现的机制可解释性技术对比。任务包含两个赛道:电路定位,评估识别因果影响组件及交互的方法;因果变量定位,评估将激活映射为可解释特征的方法。三支团队共使用八种方法,在电路定位中通过集成与正则化策略取得显著提升;一支团队采用两种方法,在因果变量定位中利用低维非线性投影实现显著进步。MIB 排行榜持续开放,鼓励后续研究在此标准框架下持续推进机制可解释性评估。

原文摘要 · Abstract (English)

Mechanistic interpretability (MI) seeks to uncover how language models (LMs) implement specific behaviors, yet measuring progress in MI remains challenging. The recently released Mechanistic Interpretability Benchmark (MIB; Mueller et al., 2025) provides a standardized framework for evaluating circuit and causal variable localization. Building on this foundation, the BlackboxNLP 2025 Shared Task extends MIB into a community-wide reproducible comparison of MI techniques. The shared task features two tracks: circuit localization, which assesses methods that identify causally influential components and interactions driving model behavior, and causal variable localization, which evaluates approaches that map activations into interpretable features. With three teams spanning eight different methods, participants achieved notable gains in circuit localization using ensemble and regularization strategies for circuit discovery. With one team spanning two methods, participants achieved significant gains in causal variable localization using low-dimensional and non-linear projections to featurize activation vectors. The MIB leaderboard remains open; we encourage continued work in this standard evaluation framework to measure progress in MI research going forward.

可解释性语言模型基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。