为心电图领域定制大模型,本地部署更安全且性能接近商用模型。
ECG-LLM -- training and evaluation of domain-specific large language models for electrocardiography
- 用心电图文献微调开源大模型,提升专业问答能力。
- 微调版Llama 3.1 70B在多项评测中表现优异,仅次于Claude 3.7。
- 适合关注医疗隐私与本地化部署的临床研究者使用。
领域适配的开源大语言模型(LLM)在医疗应用中前景广阔,如可查询的知识库和多模态助手,其关键优势在于支持本地部署以保障隐私。然而,最优适配策略、评估方法及其相对于通用大模型的性能仍不明确。我们针对心血管医学中的重要领域——心电图,通过在领域文献上微调开源模型,并构建多层次评估框架,对比微调模型、检索增强生成(RAG)与代表性通用模型Claude Sonnet 3.7。微调后的Llama 3.1 70B在多项选择题和自动文本指标上表现优越,在大模型评判测试中排名第二。人类专家评估显示,对于复杂查询,Claude 3.7和RAG方法更受青睐。微调模型在几乎所有评估模式中均显著优于基线模型。研究揭示了评估方法间的显著性能差异,凸显评估复杂性。尽管如此,通过微调与RAG实现的领域专用模型,其性能可媲美专有模型,验证了隐私保护型、本地可部署临床解决方案的可行性。
原文摘要 · Abstract (English)
Domain-adapted open-weight large language models (LLMs) offer promising healthcare applications, from queryable knowledge bases to multimodal assistants, with the crucial advantage of local deployment for privacy preservation. However, optimal adaptation strategies, evaluation methodologies, and performance relative to general-purpose LLMs remain poorly characterized. We investigated these questions in electrocardiography, an important area of cardiovascular medicine, by finetuning open-weight models on domain-specific literature and implementing a multi-layered evaluation framework comparing finetuned models, retrieval-augmented generation (RAG), and Claude Sonnet 3.7 as a representative general-purpose model. Finetuned Llama 3.1 70B achieved superior performance on multiple-choice evaluations and automatic text metrics, ranking second to Claude 3.7 in LLM-as-a-judge assessments. Human expert evaluation favored Claude 3.7 and RAG approaches for complex queries. Finetuned models significantly outperformed their base counterparts across nearly all evaluation modes. Our findings reveal substantial performance heterogeneity across evaluation methodologies, underscoring assessment complexity. Nevertheless, domain-specific adaptation through finetuning and RAG achieves competitive performance with proprietary models, supporting the viability of privacy-preserving, locally deployable clinical solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。