arXiv:2511.13095cs.CL2025-11被引 1

评测大模型在话语理解上的表现,发现其在复杂语篇推理上仍存短板。

BeDiscovER: The Benchmark of Discourse Understanding in the Era of Reasoning Language Models

  • 构建涵盖52个数据集的多层级话语理解评测基准
  • 顶尖模型在时间推理算术任务上表现强,但文档级推理能力弱
  • 适合关注大模型话语与语义理解能力的研究者使用

我们提出BeDiscovER(推理型语言模型时代的话语理解评测基准),一个全面、更新及时的现代大模型话语知识评估体系。该基准整合了5个公开的话语任务,覆盖话语词汇、句子间及文档级三个层次,共包含52个独立数据集。任务包括广泛研究的语篇解析和时间关系抽取,也包含新挑战如话语虚词消歧(例如“just”);同时汇聚了多语言、多框架的话语关系解析与树库构建共享任务。我们对开源模型Qwen3系列、DeepSeek-R1以及前沿模型GPT-5-mini进行了评测,结果表明当前最优模型在时间推理的算术层面表现优异,但在完整文档推理及某些细微语义与话语现象(如修辞关系识别)上仍存在明显困难。

原文摘要 · Abstract (English)

We introduce BeDiscovER (Benchmark of Discourse Understanding in the Era of Reasoning Language Models), an up-to-date, comprehensive suite for evaluating the discourse-level knowledge of modern LLMs. BeDiscovER compiles 5 publicly available discourse tasks across discourse lexicon, (multi-)sentential, and documental levels, with in total 52 individual datasets. It covers both extensively studied tasks such as discourse parsing and temporal relation extraction, as well as some novel challenges such as discourse particle disambiguation (e.g., ``just''), and also aggregates a shared task on Discourse Relation Parsing and Treebanking for multilingual and multi-framework discourse relation classification. We evaluate open-source LLMs: Qwen3 series, DeepSeek-R1, and frontier model such as GPT-5-mini on BeDiscovER, and find that state-of-the-art models exhibit strong performance in arithmetic aspect of temporal reasoning, but they struggle with full document reasoning and some subtle semantic and discourse phenomena, such as rhetorical relation recognition.

话语理解大模型评测推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。