arXiv:2602.10881cs.CLcs.AI2026-02中稿 · the 22nd Conferenc…被引 2

LLM在元分析证据提取中结构失效,导致结果不可靠。

Diagnosing Structural Failures in LLM-Based Evidence Extraction for Meta-Analysis

  • 设计分层查询框架,诊断LLM在复杂关系与数值绑定上的结构性缺陷。
  • 单属性提取尚可,但涉及变量绑定与效应量关联时准确率骤降。
  • 适合关注自动化元分析可靠性的研究人员,尤其警惕长文本错误放大。

系统性综述和元分析依赖将叙事性文献转化为结构化、数值化的研究记录。尽管大语言模型(LLMs)快速发展,但其是否满足该过程的结构要求仍不明确——关键在于跨文档保持角色、方法和效应量归属的完整性,而非孤立实体识别。本文提出一种结构化诊断框架,通过逐步增加关系与数值复杂度的模式化查询,精准定位超出原子级提取的失败点。基于涵盖五个科学领域的手工标注语料库,结合统一查询集与评估协议,我们在单文档与长上下文多文档输入两种场景下,评估了两种前沿LLM。结果显示:在单属性查询中表现中等,但一旦任务涉及变量、角色、统计方法与效应量间的稳定绑定,性能急剧下降;完整元分析关联三元组的提取可靠性接近零,长上下文输入进一步加剧问题。下游聚合会放大微小上游误差,使整体统计结果不可靠。分析表明,这些局限并非源于实体识别错误,而是系统性结构崩溃,包括角色颠倒、跨分析绑定漂移、密集结果段落中的实例压缩以及数值误标,揭示当前LLMs缺乏元分析所需的结构保真度、关系绑定与数值根基。代码与数据已公开于GitHub(https://github.com/zhiyintan/LLM-Meta-Analysis)。

原文摘要 · Abstract (English)

Systematic reviews and meta-analyses rely on converting narrative articles into structured, numerically grounded study records. Despite rapid advances in large language models (LLMs), it remains unclear whether they can meet the structural requirements of this process, which hinge on preserving roles, methods, and effect-size attribution across documents rather than on recognizing isolated entities. We propose a structural, diagnostic framework that evaluates LLM-based evidence extraction as a progression of schema-constrained queries with increasing relational and numerical complexity, enabling precise identification of failure points beyond atom-level extraction. Using a manually curated corpus spanning five scientific domains, together with a unified query suite and evaluation protocol, we evaluate two state-of-the-art LLMs under both per-document and long-context, multi-document input regimes. Across domains and models, performance remains moderate for single-property queries but degrades sharply once tasks require stable binding between variables, roles, statistical methods, and effect sizes. Full meta-analytic association tuples are extracted with near-zero reliability, and long-context inputs further exacerbate these failures. Downstream aggregation amplifies even minor upstream errors, rendering corpus-level statistics unreliable. Our analysis shows that these limitations stem not from entity recognition errors, but from systematic structural breakdowns, including role reversals, cross-analysis binding drift, instance compression in dense result sections, and numeric misattribution, indicating that current LLMs lack the structural fidelity, relational binding, and numerical grounding required for automated meta-analysis. The code and data are publicly available at GitHub (https://github.com/zhiyintan/LLM-Meta-Analysis).

元分析结构推理大模型缺陷自动提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。