用人类监督框架评估大模型在多源证据整合中的表现,提升研究合成的可靠性。
Knowledge Synthesis Review Framework: Task-Level Benchmarking of LLM-Based Systems for Multi-Source Evidence Synthesis

- 将证据整合拆解为筛选、提取、分析、综合四步,分任务评测大模型表现。
- 不同模型在各任务表现各异,最高准确率达82.8%,但跨源综合仍需专家判断。
- 框架支持模型动态路由,适合科研团队用于可信、可审计的智能文献综述。
快速演进领域的证据分散于学术研究、行业报告、政策文件和媒体来源,质量、结构与目的各异,导致及时整合困难。大语言模型(LLMs)或可加速此过程,但其在不同认知任务中的可靠性尚不确定。我们提出知识合成评审框架(KSR),一个以人为中心的流程,将证据整合分解为筛选、提取、分析与综合,并针对每项任务对比基于大模型的系统与专家标准,通过持续专家验证实现任务最优模型路由。在涵盖1,893份文档的语料库中,选取244份作为基准测试集,涵盖四种来源类型,使用高一致性金标准(组间一致率92.2%,kappa=0.80)进行评估。无单一系统在所有任务上领先。Claude Sonnet 4在筛选任务中准确率最高(82.8%),GPT-5召回率最高(91.8%),但特异性较低。标题与来源提取一致性超过90%,作者与参考文献字段性能下降。解释性分析与跨源综合表现最差,仍需专家判断。对截断后文档的污染检测未发现前期接触导致结果虚高。应用于全语料库时,该流程揭示了单源合成遗漏的跨源不均衡与盲点,如劳动者福祉、中小型企业及全球南方议题。KSR提供透明、可审计、模型无关的框架,保障研究合成中大模型辅助的可信性与人类责任。
原文摘要 · Abstract (English)
Evidence in rapidly evolving fields is fragmented across academic studies, industry reports, policy documents, and media sources that differ in quality, structure, and purpose, making timely synthesis difficult. Large language models (LLMs) may accelerate this work, but their reliability across the distinct cognitive tasks of a review remains uncertain. We introduce the Knowledge Synthesis Review (KSR), a human-in-the-loop framework that decomposes evidence synthesis into screening, extraction, analysis, and synthesis, benchmarks LLM-based systems on each task against expert reference standards, and routes each task to the best-performing system under continuous expert validation. We evaluated GPT-5, Claude Sonnet 4, Gemini 2.5 Pro, and NotebookLM on a 244-document benchmark subset drawn from a 1,893-document corpus on AI and work spanning four source types, against a gold standard with high inter-rater reliability (92.2% agreement, kappa = 0.80). No system led on all tasks. Claude Sonnet 4 achieved the highest screening accuracy (82.8%) and GPT-5 the highest recall (91.8%) at the expense of lower specificity. Extraction exceeded 90% agreement for titles and sources but degraded in author and reference fields. Performance declined most in interpretive analysis and cross-source synthesis, where expert judgment remained essential. A contamination check on post-cutoff documents showed no evidence that prior exposure inflated results. Applied to the full corpus, the routed workflow surfaced cross-source asymmetries and blind spots that single-source synthesis would miss, including worker well-being, small firms, and the Global South. KSR offers a transparent, auditable, model-agnostic framework for governing LLM assistance in research synthesis while preserving human accountability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。