用自洽熵面积度量推理模型生成过程的不确定性。
Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM
- 基于模型自身输出的分词级熵,动态追踪生成过程中的不确定性变化。
- 在数学题数据筛选中,同等预算下比通过率过滤提升学生模型准确率。
- 无需额外模型或采样,计算高效且结果可解释,适合训练数据质量评估。
本文提出熵面积得分(Entropy Area Score, EAS),一种简单有效的度量方法,用于量化推理型大语言模型(LLM)在答案生成过程中的不确定性。EAS 不需要外部模型或重复采样,而是利用模型自身的分词级预测熵来捕捉生成过程中不确定性的演变。实证结果表明,EAS 在不同模型和数据集上与答案熵高度相关。在训练数据选择任务中,EAS 能有效识别高潜力样本,在相同样本预算下持续优于通过率过滤策略,显著提升学生模型在数学基准测试上的准确率。EAS 兼具高效性与可解释性,为大模型训练中的不确定性建模和数据质量评估提供了实用工具。
原文摘要 · Abstract (English)
In this work, we introduce Entropy Area Score (EAS), a simple yet effective metric to quantify uncertainty in the answer generation process of reasoning large language models (LLMs). EAS requires neither external models nor repeated sampling, it integrates token-level predictive entropy from the model itself to capture the evolution of uncertainty during generation. Empirical results show that EAS is strongly correlated with answer entropy across models and datasets. In training data selection, EAS identifies high-potential samples and consistently outperforms Pass Rate filtering under equal sample budgets, improving student model accuracy on math benchmarks. EAS is both efficient and interpretable, offering a practical tool for uncertainty modeling and data quality assessment in LLM training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。