arXiv:2506.05062cs.CL2025-06EMNLP被引 6

用辩论稿评估大模型判断力,发现其与人类差异显著

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation

  • 构建600+条标注辩论稿数据集,评估模型多层理解能力
  • 大模型在局部判断接近人类,整体行为模式差异明显
  • 前沿模型可生成类人说服性演讲,适合评测推理与表达

我们提出辩论稿评估(Debate Speech Evaluation)这一新型且具有挑战性的基准,用于评估大语言模型的判断能力。该任务需深度理解演讲在论点强度与相关性、逻辑连贯性与结构、风格与语气适宜性等多个层面的表现,涉及此前系统性评测中关注较少的认知能力。为此,我们利用超过600篇精心标注的辩论稿数据集,首次深入分析了当前最先进的大模型与人类评委在此任务上的表现差异。研究发现:尽管更大模型在某些方面能逼近个体人类判断,但在整体判断行为上存在显著差异。此外,我们还考察了前沿大模型生成有说服力、立场鲜明演讲的能力,结果表明模型在该任务上已达到人类水平。

原文摘要 · Abstract (English)

We introduce Debate Speech Evaluation as a novel and challenging benchmark for assessing LLM judges. Evaluating debate speeches requires a deep understanding of the speech at multiple levels, including argument strength and relevance, the coherence and organization of the speech, the appropriateness of its style and tone, and so on. This task involves a unique set of cognitive abilities that previously received limited attention in systematic LLM benchmarking. To explore such skills, we leverage a dataset of over 600 meticulously annotated debate speeches and present the first in-depth analysis of how state-of-the-art LLMs compare to human judges on this task. Our findings reveal a nuanced picture: while larger models can approximate individual human judgments in some respects, they differ substantially in their overall judgment behavior. We also investigate the ability of frontier LLMs to generate persuasive, opinionated speeches, showing that models may perform at a human level on this task.

大模型评估辩论分析人类对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。