让大模型在可解释与高性能间动态切换,无需重新训练。
Localist LLMs -- A Mathematical Framework for Dynamic Locality Control
- 用可调参数控制注意力机制的局部化程度,实现训练推理全程调节。
- 当惩罚项超阈值时,注意力聚焦语义块,熵低且准确率高,误差可忽略。
- 适合需要透明性与强性能并存的监管场景,如医疗、金融决策系统。
我们提出一种新框架,训练大语言模型时可连续调节内部表示,覆盖从局部主义(可解释、基于规则)到分布式(泛化好、高效)的全谱系。核心创新是引入一个可调的“局部性旋钮”,在不重训练的情况下,动态控制训练和推理过程中的局部化程度。通过注意力机制的分组稀疏性惩罚、信息论锚点设计和动态规则注入实现。我们提供了严格的数学证明,明确给出了注意力集中于语义相关区块的阈值条件,附带注意力熵和指针保真度的指数级上界。具体而言,当分组稀疏性惩罚超过特定阈值时,模型注意力可精确集中在语义块上,实现低熵、高保真,且误差可忽略。该框架使从业者可在可解释模式与高性能模式间连续切换,适用于对透明性与能力均有要求的受监管领域。
原文摘要 · Abstract (English)
We present a novel framework for training large language models with continuously adjustable internal representations that span the full spectrum from localist (interpretable, rule-based) to distributed (generalizable, efficient) encodings. The key innovation is a locality dial, a tunable parameter that dynamically controls the degree of localization during both training and inference without requiring model retraining. This is achieved through group sparsity penalties on attention mechanisms, information-theoretic anchor design, and dynamic rule injection. We provide rigorous mathematical proofs establishing explicit threshold conditions under which attention provably concentrates on semantically relevant blocks, with exponential bounds on attention entropy and pointer fidelity. Specifically, we prove that when group sparsity penalties exceed certain threshold values, the model's attention mechanisms concentrate on semantically relevant blocks, achieving low entropy and high fidelity with negligible error. This framework enables practitioners to continuously interpolate between interpretable and high-performance modes, supporting applications in regulated domains requiring both transparency and capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。