评测多语言医学影像报告生成模型,发现专精语言和领域训练效果最佳。
Evaluating Vision Language Model Adaptations for Radiology Report Generation in Low-Resource Languages
- 用LLaVA框架对比通用、医学、低资源语言数据训练的视觉语言模型
- 针对意大利语、德语、西班牙语,语言特化模型显著优于其他类型
- 结合医学术语微调+低资源语言适配,适合多语言医疗AI研发
人工智能在医疗领域的应用为提升诊断与患者护理开辟了新途径。然而,在低资源语言环境下生成准确且上下文相关的放射科报告仍面临挑战。本研究构建了一个全面基准,评估指令微调的视觉-语言模型(VLMs)在意大利语、德语和西班牙语三种低资源语言中生成放射科报告的表现。基于LLaVA架构,系统性比较了使用通用数据集、领域特定数据集及低资源语言专属数据集训练的预训练模型。由于缺乏兼具医学领域与低资源语言先验知识的模型,我们分析了多种适应策略以确定最优方法。结果表明,语言特化模型在生成报告方面显著优于通用及领域特定模型,凸显语言适应的重要性。同时,经过医学术语微调的模型在所有语言中表现更优,说明领域训练的关键作用。我们还探究了温度参数对报告连贯性的影响,为最优设置提供参考。研究强调,针对语言与领域进行定制化训练对提升多语言环境下放射科报告的质量与准确性至关重要。该工作不仅深化了对VLM在医疗场景下适应能力的理解,也为未来模型调优与语言特化研究指明方向。
原文摘要 · Abstract (English)
The integration of artificial intelligence in healthcare has opened new horizons for improving medical diagnostics and patient care. However, challenges persist in developing systems capable of generating accurate and contextually relevant radiology reports, particularly in low-resource languages. In this study, we present a comprehensive benchmark to evaluate the performance of instruction-tuned Vision-Language Models (VLMs) in the specialized task of radiology report generation across three low-resource languages: Italian, German, and Spanish. Employing the LLaVA architectural framework, we conducted a systematic evaluation of pre-trained models utilizing general datasets, domain-specific datasets, and low-resource language-specific datasets. In light of the unavailability of models that possess prior knowledge of both the medical domain and low-resource languages, we analyzed various adaptations to determine the most effective approach for these contexts. The results revealed that language-specific models substantially outperformed both general and domain-specific models in generating radiology reports, emphasizing the critical role of linguistic adaptation. Additionally, models fine-tuned with medical terminology exhibited enhanced performance across all languages compared to models with generic knowledge, highlighting the importance of domain-specific training. We also explored the influence of the temperature parameter on the coherence of report generation, providing insights for optimal model settings. Our findings highlight the importance of tailored language and domain-specific training for improving the quality and accuracy of radiological reports in multilingual settings. This research not only advances our understanding of VLMs adaptability in healthcare but also points to significant avenues for future investigations into model tuning and language-specific adaptations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。