arXiv:2608.20385cs.CL2026-08

通过分析人与大模型的分歧,优化文献质量评估清单的准确性。

Using Human-LLM Disagreement to Improve Checklist-Based Quality Appraisal

论文配图:Using Human-LLM Disagreement to Improve Checklist-Based Quality Appraisal
图 1 · 摘自论文原文
  • 利用人与大模型在评估中的分歧,定位清单中模糊条目。
  • 修改模糊条目后,模型与专家判断的一致性显著提升。
  • 适合需要自动化文献评估的系统综述研究者使用。

系统综述依赖对纳入研究的质量评估,该过程耗时且易受评估清单标准模糊性影响。尽管大语言模型(LLMs)为支持此类任务提供了可能,但评估清单通常被视为固定输入,其设计如何影响与专家判断的一致性尚不明确。为此,本文研究(1)LLMs能否近似人类评估判断,以及(2)人类-模型分歧模式是否可用于识别并改进模糊的清单条目。基于《潜在轨迹研究报告指南》(GRoLTS)清单,在三个研究主题和两个清单版本下,比较了LLM生成评估与专家标注的一致性。采用条目级准确率、校正随机一致性的一致性及研究层级排序保持度进行评估。结果表明,不同条目表现差异显著,模糊与条件性条目导致最大分歧;修订这些条目后,原始一致性与校正一致性均提高。尽管条目级误判仍存在,但保留高一致性条目时,模型评分常能保持研究间的相对排序。研究显示,可靠的LLM辅助评估不仅取决于模型选择,更依赖清单设计。分析人类-模型分歧有助于识别问题条目,并支持研究合成流程的迭代优化。

原文摘要 · Abstract (English)

Systematic reviews rely on quality appraisal of included studies, a process that is time-consuming and sensitive to ambiguity in checklist criteria. Although large language models (LLMs) offer opportunities to support these tasks, appraisal checklists are typically treated as fixed inputs, and it remains unclear how their design affects agreement with expert judgments. Therefore, we investigate (1) whether LLMs can approximate human judgments in checklist-based appraisal and (2) whether patterns of human-LLM disagreement can be used to identify and improve ambiguous checklist items. Using the Guidelines for Reporting on Latent Trajectory Studies (GRoLTS) checklist, we compare LLM-generated assessments with expert annotations across three research topics and two checklist versions. Agreement is assessed using item-level accuracy, chance-corrected agreement, and preservation of study-level rank ordering. We find that performance varies substantially across checklist items, with ambiguous and conditional criteria producing the greatest disagreement. Revising these items improves both raw and chance-corrected agreement. Although item-level misclassifications persist, LLM-generated scores often preserve the relative ranking of studies when high-agreement items are retained. These results indicate that reliable LLM-assisted appraisal depends not only on model choice but also on checklist design. The findings suggest that analyzing human-LLM disagreement can help identify problematic checklist items and support the iterative improvement of research synthesis workflows.

质量评估大模型应用系统综述清单优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。