arXiv:2510.17013cs.CL2025-10被引 3

构建多语言对话追踪基准,测试模型跨段落隐含信息理解能力

DiscoTrack: A Multilingual LLM Benchmark for Discourse Tracking

  • 设计12语言4层级任务,涵盖指代消解与语篇关系推理
  • 现有顶尖模型在跨句信息整合上仍表现不佳,平均准确率不足60%
  • 适合研究多语言语篇理解、对话系统与高级推理的学者使用

当前大语言模型评测主要聚焦于单句层面的显性信息提取(如问答、摘要),而缺乏针对跨文档隐含信息与语用推理的挑战性多语言基准。为此,我们提出DiscoTrack,一个覆盖12种语言、包含四个语篇理解层级(显著性识别、实体追踪、语篇关系、桥接推理)的评测基准。评估表明,这些任务对现有最先进的模型仍具挑战性,即使在最佳模型上,跨段落信息整合的平均准确率也低于60%。该基准推动了对多语言语篇连贯性与深层推理能力的系统性评测。

原文摘要 · Abstract (English)

Recent LLM benchmarks have tested models on a range of phenomena, but are still focused primarily on natural language understanding for extraction of explicit information, such as QA or summarization, with responses often targeting information from individual sentences. We are still lacking more challenging, and importantly also multilingual, benchmarks focusing on implicit information and pragmatic inferences across larger documents in the context of discourse tracking: integrating and aggregating information across sentences, paragraphs and multiple speaker utterances. To this end, we present DiscoTrack, an LLM benchmark targeting a range of tasks across 12 languages and four levels of discourse understanding: salience recognition, entity tracking, discourse relations and bridging inference. Our evaluation shows that these tasks remain challenging, even for state-of-the-art models.

多语言语篇理解推理评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。