首个端到端乳腺钼靶报告生成框架,用大模型实现精准临床描述。
AMRG: Extend Vision Language Models for Automatic Mammography Report Generation
- 基于医域专用大模型与轻量微调,实现多视角图像理解与报告生成。
- 在DMID数据集上达到ROUGE-L 0.5691、BI-RADS准确率0.5582的先进水平。
- 适合医学影像研究者和临床AI开发者,推动放射科自动化报告发展。
乳腺钼靶报告生成是医疗AI中关键但研究不足的任务,面临多视角图像推理、高分辨率视觉线索和非结构化放射学语言等挑战。本文提出AMRG(自动乳腺钼靶报告生成),首个基于大视觉语言模型(VLM)的端到端生成框架。基于领域专用的MedGemma-4B-it模型,采用低秩适应(LoRA)参数高效微调策略,在计算开销极小的情况下完成轻量化适配。在公开数据集DMID上训练与评估,该工作建立了首个可复现的乳腺钼靶报告生成基准,填补了多模态临床AI的长期空白。系统探索了LoRA超参数配置,并在统一调优协议下对比多个VLM骨干网络,涵盖领域特异与通用模型。框架在语言生成与临床指标上表现优异,获得ROUGE-L 0.5691、METEOR 0.6152、CIDEr 0.5818以及BI-RADS准确率0.5582。定性分析显示诊断一致性提升,幻觉减少。AMRG为放射科报告生成提供可扩展、可适配的基础,推动多模态医学AI未来发展。
原文摘要 · Abstract (English)
Mammography report generation is a critical yet underexplored task in medical AI, characterized by challenges such as multiview image reasoning, high-resolution visual cues, and unstructured radiologic language. In this work, we introduce AMRG (Automatic Mammography Report Generation), the first end-to-end framework for generating narrative mammography reports using large vision-language models (VLMs). Building upon MedGemma-4B-it-a domain-specialized, instruction-tuned VLM-we employ a parameter-efficient fine-tuning (PEFT) strategy via Low-Rank Adaptation (LoRA), enabling lightweight adaptation with minimal computational overhead. We train and evaluate AMRG on DMID, a publicly available dataset of paired high-resolution mammograms and diagnostic reports. This work establishes the first reproducible benchmark for mammography report generation, addressing a longstanding gap in multimodal clinical AI. We systematically explore LoRA hyperparameter configurations and conduct comparative experiments across multiple VLM backbones, including both domain-specific and general-purpose models under a unified tuning protocol. Our framework demonstrates strong performance across both language generation and clinical metrics, achieving a ROUGE-L score of 0.5691, METEOR of 0.6152, CIDEr of 0.5818, and BI-RADS accuracy of 0.5582. Qualitative analysis further highlights improved diagnostic consistency and reduced hallucinations. AMRG offers a scalable and adaptable foundation for radiology report generation and paves the way for future research in multimodal medical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。