改进语言生成不确定性评估方法,减少虚假结果判断偏差。
Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language Generation
- 用多种评分方式替代单一正确性判断,提升评估鲁棒性
- 通过大模型评委集成降低评估偏差,效果更稳定
- 引入结构化任务和扰动检测,提供可控风险指标
幻觉是大型语言模型(LLMs)可靠性的重要障碍。近期研究指出,一类特定幻觉——虚构(confabulations)——源于模型生成时的预测不确定性。为检测此类虚构,已开发多种自然语言生成(NLG)中的不确定性估计(UE)方法。这些方法通常通过相关性评估:将不确定性估计与生成文本正确性关联,以问答(QA)数据集为标准基准。然而,常用近似正确性函数之间存在显著分歧,导致对不确定性估计方法的排序不一致,从而人为抬高其表现。本文提出使用多种替代风险指标进行风险相关性实验,增强对NLG中UE算法评估的稳健性。针对QA任务,我们证明对多个大模型作为评判者(LLM-as-a-judge)进行边际化处理可有效降低评估偏差。此外,探索结构化任务、分布外(OOD)及扰动检测任务,提供更可靠且可控的风险指标。最后,提出采用不确定性估计方法的埃洛(Elo)评级,实现多评估场景下的客观综合评价。
原文摘要 · Abstract (English)
Hallucinations are a common issue that undermine the reliability of large language models (LLMs). Recent studies have identified a specific subset of hallucinations, known as confabulations, which arise due to predictive uncertainty of LLMs. To detect confabulations, various methods for estimating predictive uncertainty in natural language generation (NLG) have been developed. These methods are typically evaluated by correlating uncertainty estimates with the correctness of generated text, with question-answering (QA) datasets serving as the standard benchmark. However, commonly used approximate correctness functions have substantial disagreement between each other and, consequently, in the ranking of the uncertainty estimation methods. This allows one to inflate the apparent performance of uncertainty estimation methods. We propose using several alternative risk indicators for risk correlation experiments that improve robustness of empirical assessment of UE algorithms for NLG. For QA tasks, we show that marginalizing over multiple LLM-as-a-judge variants leads to reducing the evaluation biases. Furthermore, we explore structured tasks as well as out of distribution and perturbation detection tasks which provide robust and controllable risk indicators. Finally, we propose to use an Elo rating of uncertainty estimation methods to give an objective summarization over extensive evaluation settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。