用多选题标准答案评估生成模型置信度,更准更省事。
MCQA-Eval: Efficient Confidence Evaluation in NLG with Gold-Standard Correctness Labels
- 基于多选题数据集的真值标签,不用复杂判断规则
- 在多个大模型上验证,评估结果更可靠
- 适合研究置信度评估或需高可信生成的场景
大型语言模型(LLMs)在医疗、法律等关键领域需要可靠的置信度估计,但现有评估框架依赖于易出错、成本高且可能引入系统偏差的正确性判定函数。我们提出MCQA-Eval,一种基于多选题数据集金标准正确标签的自然语言生成置信度评估框架,摆脱对显式正确性函数的依赖。该方法可统一评估基于内部状态的白盒(如对数概率)与基于一致性的黑盒置信度方法,在多个大模型和主流问答数据集上的实验表明,相比现有方法,MCQA-Eval能更高效、更可靠地评估置信度估计性能。
原文摘要 · Abstract (English)
Large Language Models (LLMs) require robust confidence estimation, particularly in critical domains like healthcare and law where unreliable outputs can lead to significant consequences. Despite much recent work in confidence estimation, current evaluation frameworks rely on correctness functions -- various heuristics that are often noisy, expensive, and possibly introduce systematic biases. These methodological weaknesses tend to distort evaluation metrics and thus the comparative ranking of confidence measures. We introduce MCQA-Eval, an evaluation framework for assessing confidence measures in Natural Language Generation (NLG) that eliminates dependence on an explicit correctness function by leveraging gold-standard correctness labels from multiple-choice datasets. MCQA-Eval enables systematic comparison of both internal state-based white-box (e.g. logit-based) and consistency-based black-box confidence measures, providing a unified evaluation methodology across different approaches. Through extensive experiments on multiple LLMs and widely used QA datasets, we report that MCQA-Eval provides efficient and more reliable assessments of confidence estimation methods than existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。