用物理中的重整化方法,让模型解释更可靠。
Towards Worst-Case Guarantees with Scale-Aware Interpretability
- 借鉴物理重整化理论,追踪特征在多尺度下的组合方式
- 提供对细微结构影响的量化边界,防止误判噪声为信号
- 适合关注模型可靠性与安全性的研究人员
神经网络按自然数据的分层多尺度结构组织信息。解释模型内部的方法也应具备尺度感知能力,明确追踪特征在不同分辨率下的组合,并保证对被判定为无关噪声的细粒度结构的影响有严格限制。我们提出,物理学中的重整化框架可满足这一需求,提供克服现有方法局限的技术工具。此外,相关领域的研究已成熟到足以将分散的研究成果整合为实用、基于理论的工具。为在人工智能安全背景下结合这些进展,我们提出一个统一的研究议程——尺度感知解释性,旨在发展具有鲁棒性和忠实性保障的正式机制与可解释性工具,其基础来自统计物理。
原文摘要 · Abstract (English)
Neural networks organize information according to the hierarchical, multi-scale structure of natural data. Methods to interpret model internals should be similarly scale-aware, explicitly tracking how features compose across resolutions and guaranteeing bounds on the influence of fine-grained structure that is discarded as irrelevant noise. We posit that the renormalisation framework from physics can meet this need by offering technical tools that can overcome limitations of current methods. Moreover, relevant work from adjacent fields has now matured to a point where scattered research threads can be synthesized into practical, theory-informed tools. To combine these threads in an AI safety context, we propose a unifying research agenda -- \emph{scale-aware interpretability} -- to develop formal machinery and interpretability tools that have robustness and faithfulness properties supported by statistical physics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。