arXiv:2608.30893cs.CLcs.IR2026-08

构建心电图领域语言模型评测基准,证明小模型经微调可媲美大模型。

ECGQuest: Benchmarking and Fine-Tuning Language Models for Electrocardiography

论文配图:ECGQuest: Benchmarking and Fine-Tuning Language Models for Electrocardiography
图 1 · 摘自论文原文
  • 基于23篇文献生成10,904个真/假问题,构建心电图上下文知识评测集
  • 微调后小模型最高达76.3%准确率,接近商用大模型表现
  • 适合研究医疗AI、心电图智能诊断与参数高效微调的学者使用

心电图解读需掌握心脏病学、电生理学、临床诊断、心电波形、信号采集与仪器等综合知识。现有语言模型评测主要聚焦通用医学知识或单个心电图信号/图像的解读,而非心电图解读所需的上下文知识。我们开发了ECGQuest,一个基于文献的心电图专用语言模型评测与微调资源。基于GPT-4o的流水线从2003–2025年共23篇心电图参考文献及Computing in Cardiology会议论文中生成问题,最终数据集包含10,904个唯一真/假问题及其否定形式(共21,808个问答对)。我们在零样本设置下评估了3个商用模型和20个开源模型。5个7–14B参数的开源模型采用低秩适配(LoRA)进行微调,同时纳入BERT与BiomedBERT作为监督编码器基线。泛化能力在转换为二元真/假问题的MedMCQA与MedQA心电图子集上评估,使用官方答案键。零样本准确率范围为49.5%至74.4%,其中GPT-5表现最佳。通用模型优于专科模型,部分模型存在显著真/假偏差,编码器基线表现接近随机。微调使所有开源模型准确率提升6.5–14.1%。微调后的DeepSeek-R1-Distill-Qwen-14B达到76.3%准确率,五模型投票集成达78.5%。在MedMCQA与MedQA上,微调主要提升弱模型或类别偏差模型,未一致改善强基线模型。ECGQuest提供了可复现的心电图上下文知识评测基准,表明参数高效的微调可使小型语言模型与大型商业模型竞争。

原文摘要 · Abstract (English)

Electrocardiogram (ECG) interpretation requires knowledge of cardiology, electrophysiology, clinical diagnosis, ECG waveforms, signal acquisition, and instrumentation. Existing language-model benchmarks, however, primarily assess broad medical knowledge or interpretation of individual ECG signals and images rather than the broader contextual knowledge required for ECG interpretation. We developed ECGQuest, a literature-grounded resource for evaluating and fine-tuning ECG-specific language models. A GPT-4o-based pipeline generated questions from 23 ECG references and Computing in Cardiology proceedings from 2003-2025. The final dataset contains 10,904 unique True/False questions paired with their negated forms (21,808 Q&A pairs). We evaluated three commercial and 20 open-source language models on a held-out test set in a zero-shot setting. Five open-source models with 7-14B parameters were fine-tuned using Low-Rank Adaptation, with BERT and BiomedBERT included as supervised encoder baselines. Generalization was assessed on ECG-related subsets of MedMCQA and MedQA converted to binary True/False questions using official answer keys. Zero-shot accuracy on ECGQuest ranged from 49.5% to 74.4%, with GPT-5 performing best. General-purpose models outperformed medically specialized models, several models showed strong True/False bias, and encoder baselines performed near chance. Fine-tuning improved all open-source models by 6.5-14.1%. Fine-tuned DeepSeek-R1-Distill-Qwen-14B reached 76.3% accuracy, while a five-model voting ensemble reached 78.5%. On MedMCQA and MedQA, fine-tuning mainly benefited weaker or class-biased models and did not consistently improve strong base models. ECGQuest provides a reproducible benchmark for contextual ECG knowledge and shows that parameter-efficient fine-tuning can make smaller language models competitive with substantially larger commercial models.

心电图语言模型医疗AI微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。