arXiv:2502.18993cs.CLcs.DB2025-02EMNLP被引 21

构建首个跨文档多实体问答基准,揭示大模型整合分散信息的短板

MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answering

  • 设计4780个问题,按三类八型系统化覆盖多实体推理场景
  • 顶尖模型如GPT-4在该基准上仅达59%准确率,暴露信息整合缺陷
  • 强调实体属性准确性,用EA-F1评估细粒度实体关联正确性,适合评测与改进

多实体问答(MEQA)对大语言模型(LLM)和检索增强生成(RAG)系统构成重大挑战,常难以整合来自不同文档的分散信息。现有方法虽擅长单文档理解,但在跨文档聚合方面表现不佳,尤其在处理如“ACM院士在各学科领域的分布情况”这类密集实体的问题时,需从异构来源(如维基百科页面)中整合以实体为中心的洞察。为填补这一空白,我们提出MEBench,一个新型的多文档、多实体基准,用于系统评估LLM在检索、整合和推理碎片化信息方面的能力。该基准包含4,780个问题,系统分为三大类别、八种类型,全面覆盖真实世界的多实体推理场景。我们在GPT-4、Llama-3等先进模型及RAG流水线上的实验显示,即使最先进的模型在MEBench上也仅达到59%准确率。基准强调信息提取的完整性和事实精确性,采用实体属性F1(EA-F1)指标对实体级正确性与属性归属有效性进行细粒度评估。MEBench不仅揭示了当前LLM框架的系统性弱点,也为发展更鲁棒、具实体感知能力的问答架构提供了基础。

原文摘要 · Abstract (English)

Multi-entity question answering (MEQA) represents significant challenges for large language models (LLM) and retrieval-augmented generation (RAG) systems, which frequently struggle to consolidate scattered information across diverse documents. While existing methods excel at single-document comprehension, they often struggle with cross-document aggregation, particularly when resolving entity-dense questions like "What is the distribution of ACM Fellows among various fields of study?", which require integrating entity-centric insights from heterogeneous sources (e.g., Wikipedia pages). To address this gap, we introduce MEBench, a novel multi-document, multi-entity benchmark designed to systematically evaluate LLMs' capacity to retrieve, consolidate, and reason over fragmented information. Our benchmark comprises 4,780 questions which are systematically categorized into three primary categories, further divided into eight distinct types, ensuring broad coverage of real-world multi-entity reasoning scenarios. Our experiments on state-of-the-art LLMs (e.g., GPT-4, Llama-3) and RAG pipelines reveal critical limitations: even advanced models achieve only 59% accuracy on MEBench. Our benchmark emphasizes the importance of completeness and factual precision of information extraction in MEQA tasks, using Entity-Attributed F1 (EA-F1) metric for granular evaluation of entity-level correctness and attribution validity. MEBench not only highlights systemic weaknesses in current LLM frameworks but also provides a foundation for advancing robust, entity-aware QA architectures.

多实体问答大模型评测信息整合基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。