arXiv:2506.06605cs.CLcs.AI2025-06ACL被引 16

首个可验证医学文本生成框架,提升医疗问答引用质量。

MedCite: Can Language Models Generate Verifiable Text for Medicine?

  • 提出多轮检索-引用方法,生成高质量医学引用。
  • 相比强基线模型,引用精确率与召回率显著提升。
  • 评估结果与专家标注高度一致,适合临床应用研究者。

现有基于大语言模型的医疗问答系统缺乏引用生成与评估能力,限制了实际应用。本文提出首个端到端框架MedCite,支持医学任务中引用生成的设计与评估。同时引入一种新型多轮检索-引用方法,生成高质量引用。评估揭示了医学引用生成的挑战与机遇,并识别出对最终引用质量有显著影响的关键设计选择。所提方法在引用精确率和召回率上优于强基线模型,且评估结果与专业专家标注结果高度相关。

原文摘要 · Abstract (English)

Existing LLM-based medical question-answering systems lack citation generation and evaluation capabilities, raising concerns about their adoption in practice. In this work, we introduce \name, the first end-to-end framework that facilitates the design and evaluation of citation generation with LLMs for medical tasks. Meanwhile, we introduce a novel multi-pass retrieval-citation method that generates high-quality citations. Our evaluation highlights the challenges and opportunities of citation generation for medical tasks, while identifying important design choices that have a significant impact on the final citation quality. Our proposed method achieves superior citation precision and recall improvements compared to strong baseline methods, and we show that evaluation results correlate well with annotation results from professional experts.

医学AI引用生成大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。