arXiv:2411.13760cs.LGcs.CL2024-11被引 3

提出新评估框架,解决大模型在模糊任务中的多正确答案问题。

A Framework for Evaluating LLMs Under Task Indeterminacy

  • 分离任务定义、人工评分与模型输出的关系,重构评估流程
  • 实验发现传统方法低估模型真实性能,误差可达显著水平
  • 提供带误差修正的性能区间估计法,适合复杂任务研究者

大型语言模型(LLM)评估通常假设每个测试项只有一个正确答案(黄金标签)。然而,某些任务存在歧义或模糊性——信息不足导致无法唯一确定解释,或难以明确界定判断边界。这两种情况都会引发任务不确定性:评估语料中部分项目存在多个合理回答。本文提出一种在任务不确定性下评估LLM的框架,解耦评估流程中任务设定、人工评分与模型输出之间的关系。通过合成实验表明,基于‘黄金标签’假设的评估会低估模型真实表现。我们还提出一种方法,在仅掌握部分不确定项信息的前提下,估算误差调整后的性能区间。最后,讨论了该工作对研究社区的启示。

原文摘要 · Abstract (English)

Large language model (LLM) evaluations often assume there is a single correct response -- a gold label -- for each item in the evaluation corpus. However, some tasks can be ambiguous -- i.e., they provide insufficient information to identify a unique interpretation -- or vague -- i.e., they do not clearly indicate where to draw the line when making a determination. Both ambiguity and vagueness can cause task indeterminacy -- the condition where some items in the evaluation corpus have more than one correct response. In this paper, we develop a framework for evaluating LLMs under task indeterminacy. Our framework disentangles the relationships between task specification, human ratings, and LLM responses in the LLM evaluation pipeline. Using our framework, we conduct a synthetic experiment showing that evaluations that use the "gold label" assumption underestimate the true performance. We also provide a method for estimating an error-adjusted performance interval given partial knowledge about indeterminate items in the evaluation corpus. We conclude by outlining implications of our work for the research community.

大模型评估任务不确定性性能估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。