arXiv:2603.19460cs.LGcs.CG2026-03ACL

通过几何学习提升大模型可解释性,让推理路径更透明。

GeoLAN: Geometric Learning of Latent Explanatory Directions in Large Language Models

  • 将词元表示视为几何轨迹,设计可微正则化项增强多样性。
  • 中等规模模型在保持准确率前提下,几何指标改善且公平性偏差降低。
  • 适合关注模型可解释性与内部机制分析的研究者。

大型语言模型虽表现优异,但缺乏透明性。本文提出GeoLAN训练框架,将词元表示视为几何轨迹,并借鉴凯克亚猜想相关进展引入粘附性条件。设计了两种可微正则化器:Katz-Tao Convex Wolff(KT-CW)和Katz-Tao Attention(KT-Attn),分别促进各向同性和注意力多样性。在Gemma-3(1B、4B、12B)和Llama-3-8B上的实验表明,GeoLAN在多数情况下维持任务准确率的同时,提升几何指标并减少特定公平性偏差,该优势在中等规模模型上最为显著。研究揭示了几何精度与性能之间的尺度依赖权衡,表明几何感知训练是提升机制可解释性的有前景方向。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate strong performance, but they often lack transparency. We introduce GeoLAN, a training framework that treats token representations as geometric trajectories and applies stickiness conditions inspired by recent developments related to the Kakeya Conjecture. We have developed two differentiable regularizers, Katz-Tao Convex Wolff (KT-CW) and Katz-Tao Attention (KT-Attn), that promote isotropy and encourage diverse attention. Our experiments with Gemma-3 (1B, 4B, 12B) and Llama-3-8B show that GeoLAN frequently maintains task accuracy while improving geometric metrics and reducing certain fairness biases. These benefits are most significant in mid-sized models. Our findings reveal scale-dependent trade-offs between geometric precision and performance, suggesting that geometry-aware training is a promising approach to enhance mechanistic interpretability.

可解释性几何学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。