arXiv:2608.19908cs.ITcs.LG2026-08

一种简单架构,高效估算大字符集概率

A Layered Simplex Architecture for Large Alphabets

  • 通过坐标乘积重归一化构建贝叶斯估计器,仅需调节深度参数
  • 悔恨值可显式计算,且在真实文本与合成数据上表现优于经典方法
  • 揭示了符号发现速率与数据量、词表大小的缩放规律,适合语言建模研究者

在对数损失下对大字符集进行概率估计是一个经典问题,已有如Good-Turing估计等著名方法。本文提出一种新的贝叶斯估计器,具有四大特性:其一,构造极其简单——将独立均匀采样的概率单纯形坐标逐点相乘后重新归一化,仅需调整深度这一结构参数,通过对深度平均可消除调参需求;其二,所得混合模型的悔恨值(相对于已知信源编码的额外码长)具有显式且可高效计算的表达式;其三,尽管结构简单且无调优常数,该估计器在多种合成与真实文本基准测试中表现优异,媲美甚至超越更复杂的专用方法,包括Good-Turing;其四,其悔恨值的可解析性使我们得以识别数据量、词表大小与深度之间的缩放律。对于指数大于1的Zipf分布,在样本仅揭示极小部分词表时,悔恨值可读作发现符号集合的描述长度,每比特描述仅需约1比特编码,外加每个符号的额外开销。因此,数据指数即为新符号被发现的速率。

原文摘要 · Abstract (English)

Probability estimation over large alphabets under log loss is a well-studied problem, with celebrated methods such as the Good-Turing estimator. We introduce and study a new Bayesian estimator with four notable properties. First, its construction is exceptionally simple: multiply independent uniform draws from the probability simplex coordinate-wise and renormalize. Depth is the only structural parameter, and averaging over depths eliminates the need to tune it. Second, the regret of the resulting mixture, the excess code length it pays relative to a code that knows the source, admits an explicit and efficiently computable expression. Third, despite its simplicity and lack of tuned constants, the estimator is competitive across a diverse set of synthetic and real-text benchmarks with substantially more specialized methods, including Good-Turing. Fourth, the tractability of its regret allows us to identify scaling laws in data, alphabet size, and depth. For Zipf targets with exponent above one, the regret has a simple reading as long as the sample reveals only a small fraction of the alphabet. It closely matches the description length of the set of discovered symbols, at one bit of code per bit of description, plus a further cost per symbol. The data exponent is therefore the rate at which new symbols are discovered.

概率估计大词表贝叶斯方法缩放律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。