arXiv:2510.10762cs.CLstat.AP2025-10

用大模型自动评估论文方法学,发现能识别明确信息但难处理复杂推理。

Large Language Models for Full-Text Methods Assessment: A Case Study on Mediation Analysis

  • 以中介分析为例,对比大模型与专家对180篇全文的判断
  • 模型在简单标准上接近人类准确率,复杂推断任务差15%以上
  • 适合用于快速筛查显性方法信息,需专家复核深层逻辑

系统评价对整合科学证据至关重要,但提取详细方法学信息仍耗时费力。本文以因果中介分析为典型方法领域,将先进大语言模型(LLMs)与专家评审者在180篇全文上进行对比。模型表现与人类判断高度相关(准确率相关系数0.71;F1相关系数0.97),在明确陈述的方法学标准上达到近人类水平。但在需要复杂推断的任务中,准确率显著下降,最差时落后专家15%。常见错误源于对关键词的表面误读,如将“纵向”或“敏感性”误认为方法严谨的直接证据,导致系统性误判。文档越长,模型准确率越低,但发表年份无显著影响。研究提示:当前大模型擅长识别显性方法特征,但在复杂语境下仍需人工审校。结合自动化提取与定向专家审查,可提升多领域证据合成的效率与严谨性。

原文摘要 · Abstract (English)

Systematic reviews are crucial for synthesizing scientific evidence but remain labor-intensive, especially when extracting detailed methodological information. Large language models (LLMs) offer potential for automating methodological assessments, promising to transform evidence synthesis. Here, using causal mediation analysis as a representative methodological domain, we benchmarked state-of-the-art LLMs against expert human reviewers across 180 full-text scientific articles. Model performance closely correlated with human judgments (accuracy correlation 0.71; F1 correlation 0.97), achieving near-human accuracy on straightforward, explicitly stated methodological criteria. However, accuracy sharply declined on complex, inference-intensive assessments, lagging expert reviewers by up to 15%. Errors commonly resulted from superficial linguistic cues -- for instance, models frequently misinterpreted keywords like "longitudinal" or "sensitivity" as automatic evidence of rigorous methodological approache, leading to systematic misclassifications. Longer documents yielded lower model accuracy, whereas publication year showed no significant effect. Our findings highlight an important pattern for practitioners using LLMs for methods review and synthesis from full texts: current LLMs excel at identifying explicit methodological features but require human oversight for nuanced interpretations. Integrating automated information extraction with targeted expert review thus provides a promising approach to enhance efficiency and methodological rigor in evidence synthesis across diverse scientific fields.

大模型评估方法学审查证据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。