用多模态大模型统一生成多器官多语言超声报告
Ultrasound Report Generation with Multimodal Large Language Models for Standardized Texts
- 分块式多语言训练,对齐影像与标准化文本片段
- 在多器官多语言数据上实现15%的CIDEr提升
- 适合临床落地,减少报告遗漏或错误
超声报告生成因图像差异大、操作者依赖性强且需标准化文本而具挑战性。相较于X光和CT,超声缺乏一致的数据集,自动化困难。本研究提出统一框架,支持多器官、多语言超声报告生成,通过分块式多语言训练并结合标准化报告特性,将文本片段与多样影像数据对齐。研究构建了中英双语数据集,使生成文本在不同器官和语言间保持一致性和临床准确性。通过选择性解冻视觉变换器(ViT)进行微调,进一步提升图文对齐效果。相比先前最先进方法KMVE,本方法在BLEU上提升约2%,ROUGE-L提升约3%,CIDEr提升约15%,显著降低内容缺失或错误率。该框架将多器官、多语言报告生成统一于单一可扩展系统,具备实际临床应用潜力。
原文摘要 · Abstract (English)
Ultrasound (US) report generation is a challenging task due to the variability of US images, operator dependence, and the need for standardized text. Unlike X-ray and CT, US imaging lacks consistent datasets, making automation difficult. In this study, we propose a unified framework for multi-organ and multilingual US report generation, integrating fragment-based multilingual training and leveraging the standardized nature of US reports. By aligning modular text fragments with diverse imaging data and curating a bilingual English-Chinese dataset, the method achieves consistent and clinically accurate text generation across organ sites and languages. Fine-tuning with selective unfreezing of the vision transformer (ViT) further improves text-image alignment. Compared to the previous state-of-the-art KMVE method, our approach achieves relative gains of about 2\% in BLEU scores, approximately 3\% in ROUGE-L, and about 15\% in CIDEr, while significantly reducing errors such as missing or incorrect content. By unifying multi-organ and multi-language report generation into a single, scalable framework, this work demonstrates strong potential for real-world clinical workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。