arXiv:2608.01223cs.AIcs.CE2026-08

用可调参数统一多种AI模型设计,揭示深度学习动态的非平衡特性。

Perspectives on Tsallis Statistics for Artificial Intelligence

论文配图:Perspectives on Tsallis Statistics for Artificial Intelligence
图 1 · 摘自论文原文
  • 引入参数q的广义统计框架,实现稠密与稀疏行为的连续调控。
  • 发现深度网络中重尾权重分布和梯度噪声符合q-统计特征。
  • 建议将q作为可学习的归纳偏置,适用于模型设计与优化研究者。

Tsallis统计通过单个实数参数q调节罕见与频繁事件的权重,最初用于描述具有长程关联、多分形结构和重尾波动的物理系统。该框架已成为现代人工智能中的常见组件:支撑稀疏注意力机制(如sparsemax和α-entmax)、可控探索的最大熵强化学习、鲁棒且重尾的概率模型,以及一类广义损失函数与正则化方法。本文系统梳理了其在人工智能中的应用。首先回顾数学基础:q-熵及其变分(最大熵)原理、q-指数与q-对数、q-中心极限定理、q-高斯分布及其在超统计中的动力学起源,强调对机器学习关键的性质。随后综述其在softmax推广、强化学习、序列与图神经网络、生成与概率建模、损失设计及优化中的应用,提炼出核心模式:由q控制的稠密/均匀与稀疏/尖峰行为之间的可调插值。进一步指出,深度网络中观测到的重尾权重谱和梯度噪声统计本身即为非扩展性标志,表明现代学习动态处于q-统计范畴。最后讨论方法陷阱、与信息几何及q-指数族的关系,提出q应被视为可学习的归纳偏置而非固定超参数。

原文摘要 · Abstract (English)

Tsallis statistics generalizes Boltzmann-Gibbs statistical mechanics through a single real parameter $q$ that controls the weight assigned to rare and frequent events. Originally proposed to describe physical systems with long-range correlations, multifractal geometry, and heavy-tailed fluctuations, the framework has become a recurring ingredient in modern artificial intelligence (AI): it underlies sparse attention mechanisms (\textsc{sparsemax} and $α$-\textsc{entmax}), maximum-entropy reinforcement learning with controllable exploration, robust and heavy-tailed probabilistic models, and a family of generalized loss functions and regularizers. This paper offers a structured perspective on where Tsallis statistics meets AI. We first review the mathematical core: $q$-entropy and its variational (maximum-entropy) foundation, the $q$-exponential and $q$-logarithm, the $q$-central limit theorem, $q$-Gaussian distributions, and their dynamical origin in superstatistics, emphasizing the properties that matter for machine learning. We then survey applications across softmax generalization, reinforcement learning, sequential and graph neural models, generative and probabilistic modeling, loss design, and optimization, extracting the recurring design pattern in each case: a tunable interpolation between dense/uniform and sparse/peaked behavior governed by $q$. We further argue that the heavy-tailed weight spectra and gradient-noise statistics empirically observed in deep networks are themselves nonextensive signatures, placing modern learning dynamics within the scope of $q$-statistics. Finally, we discuss methodological pitfalls, the relationship to information geometry and $q$-exponential families, and open directions, arguing that $q$ should be treated as a learnable inductive bias rather than a fixed hyperparameter.

统计力学注意力机制深度学习非平衡系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。