arXiv:2511.10658cs.CLcs.AI2025-11被引 3

15个开源大模型跨病种多语言医学报告提取,效果接近人工标注。

Evaluating Open-Weight Large Language Models for Structured Data Extraction from Narrative Medical Reports Across Multiple Use Cases and Languages

  • 对比15个开源模型在6类疾病报告上的表现,涵盖多种提示策略。
  • 小中型通用模型性能接近大型模型,少样本提示提升约13%准确率。
  • 适合医疗数据标准化、临床研究者及多语言医疗系统开发者参考。

大型语言模型(LLMs)被广泛用于从自由文本临床记录中提取结构化信息,但以往研究多集中于单一任务、有限模型和英文报告。本文在荷兰、英国和捷克的三家机构,评估了15个开源权重的LLMs在病理学与影像学报告中的表现,覆盖结直肠肝转移瘤、肝脏肿瘤、神经退行性疾病、软组织肿瘤、黑色素瘤和肉瘤共六类疾病。模型包括通用型与医学专用型,不同规模,并比较了零样本、单样本、少样本、思维链、自一致性与提示图六种提示策略。使用任务适配指标评估性能,通过共识排序聚合与线性混合效应模型量化变异性。表现最佳模型在各任务上宏平均得分接近人工标注者间的一致性水平。小至中等规模的通用模型表现媲美大型模型,而超小型与专用模型表现较差。提示图与少样本提示分别带来约13%的性能提升。任务特异性因素如复杂度与标注差异性对结果影响大于模型规模或提示策略。结果表明,开源大模型可在多疾病、多语言、多机构场景下有效提取临床报告结构化数据,为临床数据整理提供可扩展方案。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to extract structured information from free-text clinical records, but prior work often focuses on single tasks, limited models, and English-language reports. We evaluated 15 open-weight LLMs on pathology and radiology reports across six use cases, colorectal liver metastases, liver tumours, neurodegenerative diseases, soft-tissue tumours, melanomas, and sarcomas, at three institutes in the Netherlands, UK, and Czech Republic. Models included general-purpose and medical-specialised LLMs of various sizes, and six prompting strategies were compared: zero-shot, one-shot, few-shot, chain-of-thought, self-consistency, and prompt graph. Performance was assessed using task-appropriate metrics, with consensus rank aggregation and linear mixed-effects models quantifying variance. Top-ranked models achieved macro-average scores close to inter-rater agreement across tasks. Small-to-medium general-purpose models performed comparably to large models, while tiny and specialised models performed worse. Prompt graph and few-shot prompting improved performance by ~13%. Task-specific factors, including variable complexity and annotation variability, influenced results more than model size or prompting strategy. These findings show that open-weight LLMs can extract structured data from clinical reports across diseases, languages, and institutions, offering a scalable approach for clinical data curation.

医学信息提取开源模型多语言临床数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。