arXiv:2506.15830cs.CLquant-ph2025-06被引 5

用几何视角重看大模型训练,揭示优化本质。

Rethinking LLM Training through Information Geometry and Quantum Metrics

  • 以费舍尔信息度量构建参数空间几何,改进优化方向。
  • 发现曲率能解释尖锐极小点、泛化性与缩放规律。
  • 提出量子类比,为未来量子增强优化提供思路。

大型语言模型(LLMs)的优化在高维参数空间中进行,该空间具有非欧几里得结构。信息几何利用费舍尔信息度量刻画这一景观,通过自然梯度下降实现更合理的学习。尽管常不实用,但该几何视角澄清了尖锐极小点、泛化能力及观测到的缩放定律等现象。本文主张基于曲率的方法可深化对LLM训练的理解。最后,我们基于法比尼-施蒂德度量与量子费舍尔信息,推测其在量子增强系统中的高效优化潜力。

原文摘要 · Abstract (English)

Optimization in large language models (LLMs) unfolds over high-dimensional parameter spaces with non-Euclidean structure. Information geometry frames this landscape using the Fisher information metric, enabling more principled learning via natural gradient descent. Though often impractical, this geometric lens clarifies phenomena such as sharp minima, generalization, and observed scaling laws. We argue that curvature-based approaches deepen our understanding of LLM training. Finally, we speculate on quantum analogies based on the Fubini-Study metric and Quantum Fisher Information, hinting at efficient optimization in quantum-enhanced systems.

大模型训练信息几何优化理论量子计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。