arXiv:2502.03298cs.CLcs.AI2025-02被引 8

构建患者友好问答数据集,评估大模型医疗简化效果

MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters

  • 用大模型自动生成并人工校验出院小结问答对
  • 通用大模型表现优于专业医学模型,自动评估与人工一致
  • 适合医疗AI可解释性研究和患者信息辅助系统开发

尽管提升患者获取医疗文档的机会有助于改善医疗服务,但受限于健康素养差异和复杂的医学术语。大型语言模型(LLMs)可通过简化医学信息提供解决方案。然而,由于缺乏标准化评估资源,评估LLMs生成安全且面向患者的文本仍具挑战性。为此,我们构建了MeDiSumQA:一个基于MIMIC-IV出院小结,通过结合大模型问答生成与人工质量检查的自动化流程创建的数据集。利用该数据集,我们评估了多种LLMs在面向患者的问答任务上的表现。结果表明,通用型大模型频繁优于生物医学优化模型,且自动化指标与人类判断高度相关。通过在PhysioNet发布MeDiSumQA,我们旨在推动大模型发展,以增强患者理解力,最终改善临床结局。

原文摘要 · Abstract (English)

While increasing patients' access to medical documents improves medical care, this benefit is limited by varying health literacy levels and complex medical terminology. Large language models (LLMs) offer solutions by simplifying medical information. However, evaluating LLMs for safe and patient-friendly text generation is difficult due to the lack of standardized evaluation resources. To fill this gap, we developed MeDiSumQA. MeDiSumQA is a dataset created from MIMIC-IV discharge summaries through an automated pipeline combining LLM-based question-answer generation with manual quality checks. We use this dataset to evaluate various LLMs on patient-oriented question-answering. Our findings reveal that general-purpose LLMs frequently surpass biomedical-adapted models, while automated metrics correlate with human judgment. By releasing MeDiSumQA on PhysioNet, we aim to advance the development of LLMs to enhance patient understanding and ultimately improve care outcomes.

医疗AI大模型问答生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。