提出最优消融方法,更精准定位模型关键组件。
Optimal ablation for interpretability
- 基于理论最优性设计消融策略,提升重要性评估准确性。
- 在电路发现等任务中显著优于传统消融方法。
- 适合研究模型内部机制的学者与可解释性开发者。
可解释性研究常通过追踪机器学习模型中的信息流,识别对特定任务有贡献的模型组件。以往工作通过消融某个组件或模拟禁用该组件时的模型推理来量化其重要性。本文提出一种新方法——最优消融(Optimal Ablation, OA),并证明基于OA的组件重要性在理论上和实证上均优于其他消融方法。此外,基于OA的重要性度量还能有效提升多个下游可解释性任务的表现,包括电路发现、事实记忆定位和潜在表示预测。
原文摘要 · Abstract (English)
Interpretability studies often involve tracing the flow of information through machine learning models to identify specific model components that perform relevant computations for tasks of interest. Prior work quantifies the importance of a model component on a particular task by measuring the impact of performing ablation on that component, or simulating model inference with the component disabled. We propose a new method, optimal ablation (OA), and show that OA-based component importance has theoretical and empirical advantages over measuring importance via other ablation methods. We also show that OA-based component importance can benefit several downstream interpretability tasks, including circuit discovery, localization of factual recall, and latent prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。