评测大模型在长文本生成中表达不确定性的能力,发现现有模型表现不佳。
UNCLE: Benchmarking Uncertainty Expressions in Long-Form Generation
- 构建首个长短文本对齐的不确定性评测基准UNCLE
- 1000+实体覆盖5个领域,验证模型在长文本中难以正确表达不确定
- 训练方法比提示工程更有效,为未来研究提供方向
大型语言模型在长文本生成中容易产生幻觉。一种有前景的缓解方法是教会模型在知识不足时显式表达不确定性。然而,现有工作缺乏对模型在长文本生成中表达不确定性的直接、公平评估。为此,我们提出UNCLE基准,用于评估长短文本问答中的不确定性表达能力。该基准涵盖5个领域,包含超过1000个实体,每项均配有对齐的短文本与长文本问答对。我们的数据集是首个通过一致问题和标准答案实现长短文本对齐的。同时,我们设计了一套新指标来评估模型选择性表达不确定性的能力。实验表明,当前模型在长文本生成中未能恰当表达不确定性。我们进一步探索了基于提示和基于训练的方法,后者提升更显著。对长短文本不确定性表达之间对齐差距的分析,揭示了未来研究的潜在方向。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are prone to hallucination, particularly in long-form generations. A promising direction to mitigate hallucination is to teach LLMs to express uncertainty explicitly when they lack sufficient knowledge. However, existing work lacks direct and fair evaluation of LLMs' ability to express uncertainty effectively in long-form generation. To address this gap, we first introduce UNCLE, a benchmark designed to evaluate uncertainty expression in both long- and short-form question answering (QA). UNCLE covers five domains and includes more than 1,000 entities, each with paired short- and long-form QA items. Our dataset is the first to directly link short- and long-form QA through aligned questions and gold-standard answers. Along with UNCLE, we propose a suite of new metrics to assess the models' capabilities to selectively express uncertainty. We then demonstrate that current models fail to convey uncertainty appropriately in long-form generation. We further explore both prompt-based and training-based methods to improve models' performance, with the training-based methods yielding greater gains. Further analysis of alignment gaps between short- and long-form uncertainty expression highlights promising directions for future research using UNCLE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。