arXiv:2502.14366cs.CLcs.AI2025-02

用信息论优化选词,让生成文本更均衡自然。

Entropy-UID: A Method for Optimizing Information Density

  • 联合最小化熵与意外度,动态调整选词策略。
  • 在多个数据集上降低意外度和熵方差,提升平衡性。
  • 适合追求生成质量与稳定性的语言模型研究者。

平衡高效的信息流对优化语言生成模型至关重要。本文提出熵-UID(Entropy-UID)新方法,通过联合最小化熵与意外度(surprisal),实现对令牌选择的自适应调整,促进生成序列中信息分布更均匀。理论证明该方法能最优减少信息突增,同时保持流畅性与连贯性。在WikiText-2、OpenWebText和WMT等多个基准数据集上,使用信息论指标评估表明,相比标准GPT-2和其它启发式方法,熵-UID在生成文本中实现了更低的意外度与熵方差,生成结果更具平衡性与人类写作相似性。研究结果表明,利用信息论约束可有效改进自回归语言模型中的令牌选择策略。

原文摘要 · Abstract (English)

Balanced and efficient information flow is essential for optimizing language generation models. In this work, we propose Entropy-UID, a new token selection method that balances entropy and Uniform Information Density (UID) principles for enhanced efficiency of text generation. Our approach adaptively adjusts token selection by jointly minimizing entropy and surprisal, promoting more even information distribution across generated sequences. Theoretical validation demonstrates that Entropy-UID optimally reduces information spikes while maintaining fluency and coherence. The method has been evulated using information-theoretic metrics on multiple benchmark datasets, including WikiText-2, OpenWebText, and WMT. Experimental results show that Entropy-UID achieves lower surprisal and entropy variance compared to standard GPT-2 and alternative heuristics, leading to more balanced and human-like text generation. Our findings point towards the potential of leveraging information-theoretic constraints to refine token selection strategies in autoregressive language models.

语言模型信息论生成质量令牌选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。