评测大模型在超长辩论文本中的推理与论证能力,揭示当前模型的不足。
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models
- 构建包含32场超过1小时辩论的长文本数据集,每篇平均3.2万词元。
- 模型需分析8篇7分钟演讲并给出最终判断,正确率普遍偏低。
- 适合研究长上下文推理、论辩生成与人类专家对齐的学者使用。
我们提出DebateBench,一个由世界顶级辩论赛事中精选的英国议会制辩论转录稿及元数据组成的全新数据集。该数据集涵盖32场辩论,共256篇演讲,每场辩论时长超1小时,每篇输入平均达32,000个词元。所有内容均附有官方评委评分与队伍排名。本数据集旨在评估大语言模型(LLMs)在长上下文推理、论辩分析与人类专家对齐方面的能力。要在此任务上表现良好,模型必须通过上下文学习理解辩论规则与评判标准,分析多位发言人的复杂论点,并作出合理裁决。初步测试GPT-o1、GPT-4o与Claude Haiku显示,现有模型在该任务上表现不佳,凸显提升其长程推理能力的紧迫性。
原文摘要 · Abstract (English)
We introduce DebateBench, a novel dataset consisting of an extensive collection of transcripts and metadata from some of the world's most prestigious competitive debates. The dataset consists of British Parliamentary debates from prestigious debating tournaments on diverse topics, annotated with detailed speech-level scores and house rankings sourced from official adjudication data. We curate 256 speeches across 32 debates with each debate being over 1 hour long with each input being an average of 32,000 tokens. Designed to capture long-context, large-scale reasoning tasks, DebateBench provides a benchmark for evaluating modern large language models (LLMs) on their ability to engage in argumentation, deliberation, and alignment with human experts. To do well on DebateBench, the LLMs must perform in-context learning to understand the rules and evaluation criteria of the debates, then analyze 8 seven minute long speeches and reason about the arguments presented by all speakers to give the final results. Our preliminary evaluation using GPT o1, GPT-4o, and Claude Haiku, shows that LLMs struggle to perform well on DebateBench, highlighting the need to develop more sophisticated techniques for improving their performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。