arXiv:2608.15798cs.LGstat.ML2026-08

语言模型的交叉熵风险无法被一致估计,即使使用留出法也不行。

Cross-Entropy Risk Estimation for Language Models: Inconsistency Must Be Dense, and the Holdout Method Is No Exception

  • 风险估计一致性依赖于数据分布与模型权重的联合状态,样本无法揭示其尾部特性
  • 无论是否限制序列长度或模型支持范围,不一致估计在所有可定义风险的状态中都普遍存在
  • 两种解决路径:有限上下文窗口或预设阈值报告,各有代价且需重新理解估计目标

语言模型通过其留出的每标记交叉熵风险进行比较——这是规模定律所拟合的量。我们证明该风险无法被一致估计。一致性定义于一个可能世界下:即数据生成分布与实际训练模型的组合。量化模型与生成机制是必要的,因为决定风险是否可估的是模型权重诱导分布的尾部性质,而样本无法揭示这一点。每标记交叉熵风险难以估计源于一个拓扑事实:在所有可能状态中,有限风险与无限风险彼此任意接近。因此,任何估计器——不仅限于留出平均——在风险有定义的所有状态下均不一致。更糟的是,即使在限制期望序列长度或全支撑模型的情况下,不一致估计仍持续存在;而在该受限设定中,不一致发生的状态甚至更为密集。提出了两个有趣出路:第一,使用有限上下文窗口可对模型下一个标记概率进行下界处理,使其风险为有限当且仅当数据生成分布具有有限期望序列长度——为原本出于计算考虑的选择提供了新的统计依据,但该假设本身也无法被任何检验所验证;第二,仅在风险低于预先设定阈值时报告,可恢复一致性,且不影响模型选择的实际需求——但需认识到估计目标已被重新定义。

原文摘要 · Abstract (English)

Language models are compared by their held-out per-token cross-entropy risk---the quantity scaling laws are fitted to. We show that it cannot be consistently estimated. Consistency, or convergence to the estimand, is defined relative to a \emph{possible state of the world}: a pair consisting of a data-generating distribution and a model we turn out to train. Quantifying over models as well as data-generating mechanisms is essential, because what decides whether a model's risk is estimable is a tail property of the distribution its weights induce, which no sample reveals. The per-token cross-entropy risk is hard to estimate because of a topological fact: among the possible states, finite risk and infinite risk each lie arbitrarily close to every instance of the other. Consequently no estimator---not merely the holdout average---is consistent at every state at which the risk is defined. Worse, inconsistent estimation persists under both bounding the expected sequence length and restricting to full-support models; and in that restricted setting the states at which inconsistency occurs are even dense. Two interesting ways out are identified, and neither is free. Way out 1: using a bounded context window, we can floor a model's next-token probabilities, making its risk finite exactly when the data-generating distribution has finite expected sequence length---a new, statistical rationale for a choice that was made on computational grounds, though the assumption it substitutes is itself beyond the reach of any test. Way out 2: reporting the risk only when it falls below a threshold fixed in advance restores consistency, at no cost to what model selection actually requires---but we need to recognize that the goal of estimation is revised.

语言模型风险估计一致性交叉熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。