arXiv:2607.27654cs.CLcs.AI2026-07

构建多粒度事件分析基准,评估大模型跨文档事件理解能力

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

论文配图:From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models
图 1 · 摘自论文原文
  • 用自校正框架自动构建高质量事件数据集
  • 设计四类任务覆盖从单句到跨文档的事件分析
  • 揭示大模型在复杂事件推理中的关键短板

事件分析是信息抽取的核心方向,涵盖不同文档粒度下的多种以事件为中心的任务。尽管大语言模型(LLMs)在部分任务上已取得良好表现,但现有基准受限于文档粒度、任务设计和数据来源,难以全面评估其事件分析能力。为此,我们提出MiGUE-Bench,一个系统化的多粒度事件分析评估基准。为支持大规模评估,我们开发了基于LLM的自校正标注框架MiGUE-Pipeline,实现高质事件数据的可扩展获取与自动标注。基准包含四项核心任务:事件检测、关系推理、结构归纳和未来预测,分别考察模型在原子事件细节与复杂跨文档叙事中的能力。对前沿LLMs及检索增强生成(RAG)方法的广泛实验,明确了当前能力边界并识别出关键缺陷,为未来提升大模型在挑战性事件分析任务中的表现提供了重要洞见。

原文摘要 · Abstract (English)

Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved promising performance in part of these tasks individually, their capability in event analysis still lacks comprehensive understanding due to restricted document granularity, task designs, and data source of existing benchmarks. To address these limitations, we introduce MiGUE-Bench, a systematic benchmark for assessing the performance of LLMs in multi-granularity event analysis. To support large-scale evaluation, we first develop an LLM-driven self-correcting annotation framework called MiGUE-Pipeline, enabling scalable acquisition of high-quality source data of events with automatic labels. Then, we design four core tasks in our benchmark, i.e., event detection, relation reasoning, structure induction, and future prediction, to probe model competence at different levels, from atomic event details to complex cross-document narratives. Extensive experiments on state-of-the-art LLMs and retrieval-augmented generation (RAG) methods delineate the current capability boundary and identify critical deficiencies, providing insights into the future improvement of LLMs in challenging event analysis tasks.

事件分析大模型评估多粒度跨文档

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。