构建首个用于心电图诊断的多轮对话评测数据集,提升AI临床推理能力。
ECG-Expert-QA: A Benchmark for Evaluating Medical Large Language Models in Heart Disease Diagnosis
- 融合真实与合成心电图数据,设计12类诊断任务
- 包含47,211对专家验证的问答对,覆盖复杂罕见病
- 支持多轮对话,适合评估医疗大模型的临床推理能力
我们提出ECG-Expert-QA,一个用于评估心电图(ECG)解读中诊断能力的综合性多模态数据集。该数据集结合真实临床心电图数据与系统生成的合成病例,涵盖12项关键诊断任务,共47,211对专家验证的问答对。内容覆盖从基础心律识别到涉及罕见疾病和时间变化的复杂诊断等多样化临床场景。其核心创新在于支持多轮对话,推动可模拟医生-患者或医际互动的对话式医疗AI系统发展,实现对AI模型临床推理、诊断准确性和知识整合能力的更真实评估。数据集通过知识引导框架构建并经过严格质量控制,确保语言与临床一致性,是推进心电图辅助诊断AI的重要资源。它挑战模型识别细微缺血改变及在上下文丰富的场景中解读复杂心律失常的能力。为促进研究透明与协作,数据集、配套代码与提示模板已公开发布于https://github.com/Zaozzz/ECG-Expert-QA。
原文摘要 · Abstract (English)
We present ECG-Expert-QA, a comprehensive multimodal dataset for evaluating diagnostic capabilities in electrocardiogram (ECG) interpretation. It combines real-world clinical ECG data with systematically generated synthetic cases, covering 12 essential diagnostic tasks and totaling 47,211 expert-validated QA pairs. These encompass diverse clinical scenarios, from basic rhythm recognition to complex diagnoses involving rare conditions and temporal changes. A key innovation is the support for multi-turn dialogues, enabling the development of conversational medical AI systems that emulate clinician-patient or interprofessional interactions. This allows for more realistic assessment of AI models' clinical reasoning, diagnostic accuracy, and knowledge integration. Constructed through a knowledge-guided framework with strict quality control, ECG-Expert-QA ensures linguistic and clinical consistency, making it a high-quality resource for advancing AI-assisted ECG interpretation. It challenges models with tasks like identifying subtle ischemic changes and interpreting complex arrhythmias in context-rich scenarios. To promote research transparency and collaboration, the dataset, accompanying code, and prompts are publicly released at https://github.com/Zaozzz/ECG-Expert-QA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。