arXiv:2608.26131cs.CL2026-08

构建真实对话基准,评估大模型长对话能力

Evaluating Language Models in Realistic Conversational Contexts

论文配图:Evaluating Language Models in Realistic Conversational Contexts
图 1 · 摘自论文原文
  • 引入专业写作者生成的真人级长对话数据集
  • 发现现有自动评估方法与专家判断相关性低
  • 提出混合评估框架,提升与人工评分的相关性30%

随着大语言模型越来越多地用于开放域、多轮对话,以人类规模评估对话质量成为核心挑战。现有针对摘要、翻译或短文本问答的评估框架难以有效衡量人类规模对话的一致性,且其指标常依赖合成数据而非真实人类反馈。本文提出UPHELD(UPwork Human-Scale Evaluated Long Dialogues),一个大规模、带参考答案的基准,用于评估超越事实正确性的长对话能力。UPHELD包含数百段由专业编剧创作的真实人类对话,具有自然对话轮次密度,并涵盖36,000+条每轮的人类标注,覆盖超过30,000个专家生成的对话轮次。利用UPHELD,我们系统评估了经典自动评估指标和无参考的LLM作为裁判的方法,发现它们与专家判断相关性不足。基于此分析,我们开发了多裁判混合框架,将多种评估信号融合,使与人工评估的相关性提升约30%。UPHELD为人类规模对话智能提供了可靠、以人类为基础的评估基础,填补了现有大模型数据集的空白。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) are increasingly deployed to serve open-ended, multi-turn interactions, evaluating conversational quality at human scale has become a central challenge. Existing evaluation frameworks built for summarization, translation, or short-form QA tasks fall short of adequately measuring the consistency of human-scale dialogue, especially when derivation and validation of these metrics themselves often rely on synthetic rather than human sources. We fill the gap by introducing UPHELD (UPwork Human-Scale Evaluated Long Dialogues), a large, reference-full benchmark for evaluating human-scale conversational ability beyond factual correctness. UPHELD consists of hundreds of complete human-to-human dialogues authored by professional script writers, with realistic turn densities and 36,000+ per-turn human annotations across 30,000+ expert-generated dialogue turns. Using UPHELD, we systematically evaluate classical automatic metrics and reference-free LLM-as-a-judge approaches, and find them unreliable when correlated with expert human judgment. Building off this analysis, we use UPHELD to develop a Mixture-of-Judges framework that combines multiple evaluative signals and improves correlation with human assessments by approximately 30%. Overall, UPHELD provides a robust, human-grounded foundation for evaluating human-scale conversational intelligence that fills a crucial gap in the pre-existing LLM dataset landscape.

对话评估大模型评测真实对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。