推动可解释性研究可审计化,建立持续协作的评审机制
Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing

- 构建协作评审平台,持续记录方法论争议与复现结果
- 发现两篇论文对同一现象结论冲突,因方法不一致无法比较
- 适合关注AI安全、模型可信验证的研究者与政策制定者
尽管机制可解释性(MI)已揭示神经网络内部运作的重要规律,但该领域尚未建立标准化的实验审计体系。因此,其多数成果难以在医疗AI、自动驾驶等高风险场景中应用,因无法证明其有效性。近期研究证实:两篇论文对同一行为得出矛盾结论,第三项研究指出两者均部分正确,但因方法不一致而不可比。缺乏标准审计机制导致此类模糊性阻碍了高要求场景下的采纳。本文呼吁建立新型评审体系,补充传统同行评审:(1) 通过‘协作评审平台’实现持续评审,归档元科学成果与讨论(如批判、负结果、事后扩展、复现、复制和部分结果),支持随时评论与修改;(2) 将平台上提炼的良好实践转化为专家认证的指南与协议,提升审计效率;(3) 建立基于来源的审计系统,追踪主张所依赖的证据源。本文鼓励围绕该框架的必要性、设计与实施展开建设性讨论,并提供早期实例以促进对话。总体而言,我们主张对机制可解释性自身进行审计,是其在人工智能安全、产业及治理中落地的关键。
原文摘要 · Abstract (English)
While mechanistic interpretability (MI) has produced important insights into neural network internals, the field has yet to establish a standardized system to audit experiments. As such, many of its findings remain underutilized in safety-critical applications such as medical AI and autonomous systems, as stakeholders cannot certify their validity. Recent work demonstrates this concretely: two papers found conflicting conclusions for the same behavior, and a third study revealed that both were partially correct but incomparable due to methodological inconsistencies. Without standardized auditing, such ambiguities hinder adoption in high-stakes contexts requiring strong correctness guarantees. We call for the MI community to work towards developing a novel reviewing system that complements peer review via: (1) Continuous reviewing supported by a \emph{Collaborative Reviewing Platform} where meta-science results and discussions (such as critiques, negative results, post-hoc extensions, reproductions, replications, and partial results) that fit outside of papers are organized and discussed, allowing for comments and revisions to be made at any time (2) Generalizing good practices found on this platform into expert-verified guidelines and protocols to improve auditing efficiency, and (3) Source-based auditing systems that track arguments which claims depend on. This position paper encourages constructive debate over the necessity, design and implementation of such a framework, providing early concrete examples to help catalyze these dialogues. Overall, we propose that auditing MI itself is essential for its application in AI safety, industry, and governance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。