arXiv:2409.18630cs.LGcond-mat.stat-mech2024-09被引 1

用统计力学解释AI学习中的数据集中现象

Entropy, concentration, and learning: a statistical mechanics primer

  • 从统计力学出发分析模型训练中的样本集中机制
  • 揭示指数族分布与信息熵在学习中的核心作用
  • 适合对理论基础感兴趣的机器学习研究者

通过统计力学的视角,本研究从第一性原理出发,探讨了人工智能和机器学习中支撑学习过程的样本集中行为。文章强调了指数族分布、统计量以及信息论与物理学量之间的深层联系,阐明了损失最小化训练背后的理论基础。这些理论框架不仅为理解模型泛化能力提供了新视角,也为设计更高效的优化算法奠定了理论根基。

原文摘要 · Abstract (English)

Artificial intelligence models trained through loss minimization have demonstrated significant success, grounded in principles from fields like information theory and statistical physics. This work explores these established connections through the lens of statistical mechanics, starting from first-principles sample concentration behaviors that underpin AI and machine learning. Our development of statistical mechanics for modeling highlights the key role of exponential families, and quantities of statistics, physics, and information theory.

统计力学机器学习理论信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。