arXiv:2606.05436cs.AIcs.CL2026-06

AI可生成接近专家水平的医学文献摘要,但专家仍更偏好真人撰写。

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison

  • 用RAG框架结合三款大模型生成头痛领域文献摘要。
  • 专家评审显示人工摘要在正确性与临床价值上更优,但难分辨真假。
  • 揭示了人类专家重视的摘要质量维度,助力未来优化AI生成流程。

为支持循证医学和高质量患者照护,及时总结最新医学文献至关重要,但临床医生面临时间有限与文献激增的双重挑战。尽管检索增强型大语言模型(LLMs)在临床摘要方面展现出潜力,其在整合广泛科学文献方面的有效性及与专家撰写的直接对比仍缺乏充分评估。本研究构建了一个基于RAG的智能体AI框架,采用三种前沿大模型:Sonnet、GPT-4o 和 Llama 3.1。一位头痛专科医生提出13个问题(其中3个用于提示优化,10个用于评估)。来自美国和加拿大的十位头痛专家各针对一个问题撰写摘要,每题产生四份摘要(专家、Sonnet、GPT-4o、Llama)。专家在不知作者身份的情况下,依据正确性、完整性、简洁性和临床实用性,对除自己撰写的以外的摘要进行评分(1–10分),并排序偏好,判断是否为人类或AI生成。结果显示,专家更倾向于人类撰写的摘要,但有时难以区分人类与AI生成内容。研究还识别出专家看重的关键特征,可指导未来人类与AI文献摘要流程的改进。

原文摘要 · Abstract (English)

Summarizing the latest medical literature to guide clinical decision-making is essential for evidence-based medicine and high-quality patient care. Yet clinicians face increasing challenges due to limited time with patients and a rapidly growing volume of published articles. Although retrieval-augmented large language models (LLMs) have shown promise in clinical summarization, human evaluations of their effectiveness in synthesizing broader scientific literature and direct comparisons to expert-written syntheses remain scarce. We constructed a RAG-based agentic AI framework using three state-of-the-art LLMs: Sonnet, GPT-4o, and Llama 3.1. A headache specialist created 13 questions, three for prompt optimization and ten for evaluation. Ten headache specialists across the United States and Canada each wrote a summary for one question, yielding four summaries per question (expert, Sonnet, GPT-4o, and Llama). The experts, blinded to authorship, critically evaluated the summaries, excluding the topic for which they wrote a summary, based on correctness, completeness, conciseness, and clinical utility, scoring each from 1 to 10 using standardized rubrics. They also ranked the summaries by preference and indicated whether they believed each summary was written by an expert or an LLM. Our study, comparing LLM- and expert-written literature summaries evaluated by headache specialists, showed that expert-written summaries were preferred, although experts sometimes found it challenging to distinguish between human- and AI-generated summaries. We also identified key expert-valued features beyond standard evaluation metrics that can guide future refinement of both human and AI literature summarization pipelines.

医学摘要大模型临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。