通过渐进式注意力聚焦,让大模型既可解释又保持高性能。
Progressive Localisation in Localist LLMs
- 从早期分布式到晚期局部化,逐步增强注意力聚焦
- 接近基线性能的同时实现可解释的注意力模式
- 适合追求可解释性与性能平衡的研究者
本文表明,渐进式局部化——即注意力局部性从早期分布式层逐渐过渡到晚期局部化层——是构建可解释大语言模型(LLMs)的最优架构,同时保持性能。通过对在《人工超级智能心理学》数据集上微调的GPT-2进行系统实验,评估了五种局部性配置:两种均匀基线(全分布式与全局部化)及三种渐进多项式调度。研究探讨了可解释性约束是否可与自然语义结构对齐,并在网络深度上策略性应用。结果表明,结合自适应语义块划分与陡峭多项式局部化调度的渐进语义局部化方法,在保持接近基线语言建模性能的同时,提供了可解释的注意力模式。多次独立训练(不同随机种子)验证了结果的统计稳健性和高度可复现性。该方法显著优于固定窗口局部化和朴素均匀局部化约束。分析显示,通过低保真度约束维持灵活性,可保留模型容量并带来可解释性收益;在决策关键的最后几层集中局部性,而早期层保持分布式学习,能实现接近基线的注意力分布特征。这些发现表明,可解释性机制应与语义结构对齐,以实现可信AI系统的实用性能-可解释性权衡。
原文摘要 · Abstract (English)
This paper demonstrates that progressive localization, the gradual increase of attention locality from early distributed layers to late localized layers, represents the optimal architecture for creating interpretable large language models (LLMs) while preserving performance. Through systematic experimentation with GPT-2 fine-tuned on The Psychology of Artificial Superintelligence, we evaluate five locality configurations: two uniform baselines (fully distributed and fully localist) and three progressive polynomial schedules. We investigate whether interpretability constraints can be aligned with natural semantic structure while being applied strategically across network depth. We demonstrate that progressive semantic localization, combining adaptive semantic block partitioning with steep polynomial locality schedules, achieves near-baseline language modeling performance while providing interpretable attention patterns. Multiple independent training runs with different random seeds establish that results are statistically robust and highly reproducible. The approach dramatically outperforms both fixed-window localization and naive uniform locality constraints. Analysis reveals that maintaining flexibility through low-fidelity constraints preserves model capacity while providing interpretability benefits, and that steep schedules concentrating locality in decision-critical final layers while preserving distributed learning in early layers achieve near-baseline attention distribution characteristics. These findings demonstrate that interpretability mechanisms should align with semantic structure to achieve practical performance-interpretability tradeoffs for trustworthy AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。