arXiv:2503.02365cs.AIcs.CL2025-03被引 2

构建77万条心脏超声报告问答数据,助力AI辅助心病诊断。

EchoQA: A Large Collection of Instruction Tuning Data for Echocardiogram Reports

  • 基于重症监护数据库构建大规模心超报告QA数据集
  • 微调后大模型在多项指标上表现显著提升
  • 评估模型公平性,适配临床辅助诊断场景

我们引入了一个新型的问答(QA)数据集,使用来自医学信息仓储系统(MIMIC-IV)的心脏超声报告。该数据集专为提升心脏病学领域的问答系统而设计,包含771,244个问答对,涵盖广泛的心脏异常及其严重程度。我们对比了多种大语言模型(LLMs),包括开源和生物医学专用模型的零样本评估,以及闭源模型的零样本和三样本评估。结果表明,微调后的模型在各项问答指标上均有提升,验证了该数据集的价值。临床医生还对表现最佳的模型进行定性评估,以检验其回答的准确性。此外,我们进行了细粒度的公平性审计,评估模型在不同健康社会决定因素下的偏差-性能权衡。我们的目标是通过建立一个面向支持临床医生进行心脏鉴别诊断的大型语言模型基准,推动该领域发展,减轻文档负担,缓解临床医生职业倦怠,使医疗人员能更专注于患者照护。

原文摘要 · Abstract (English)

We introduce a novel question-answering (QA) dataset using echocardiogram reports sourced from the Medical Information Mart for Intensive Care database. This dataset is specifically designed to enhance QA systems in cardiology, consisting of 771,244 QA pairs addressing a wide array of cardiac abnormalities and their severity. We compare large language models (LLMs), including open-source and biomedical-specific models for zero-shot evaluation, and closed-source models for zero-shot and three-shot evaluation. Our results show that fine-tuning LLMs improves performance across various QA metrics, validating the value of our dataset. Clinicians also qualitatively evaluate the best-performing model to assess the LLM responses for correctness. Further, we conduct fine-grained fairness audits to assess the bias-performance trade-off of LLMs across various social determinants of health. Our objective is to propel the field forward by establishing a benchmark for LLM AI agents aimed at supporting clinicians with cardiac differential diagnoses, thereby reducing the documentation burden that contributes to clinician burnout and enabling healthcare professionals to focus more on patient care.

医疗AI问答系统心超报告大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。