用专家关键词提升眼底图报告生成,小样本下仍表现卓越
DREAM: Dynamic Retinal Enhancement with Adaptive Multi-modal Fusion for Expert Precision Medical Report Generation

- 双阶段融合:先对齐图像与专家关键词,再动态加权融合
- 在DeepEyeNet上达成0.241的BLEU-4,超越现有模型
- 适合医疗AI研究者与临床辅助系统开发者
自动化眼底图像医学报告生成需要结合视觉模式识别与深层临床知识。当前大型视觉语言模型在数据稀缺的专业医学领域常出现过拟合,忽略细微但关键的病理特征。为此,我们提出DREAM(动态视网膜增强自适应多模态融合),一种在数据有限条件下仍能生成高保真报告的新框架。DREAM采用独特两阶段融合机制:首先,Abstractor模块将图像与眼科医生标注的临床关键词特征映射至共享空间,以病理相关信息增强视觉表征;其次,Adaptor模块执行自适应多模态融合,通过可学习参数动态调整各模态权重,生成统一表示。为确保输出语义符合临床实际,训练中引入对比对齐模块,将融合表示与真实报告对齐。该方法在DeepEyeNet基准上达到0.241的BLEU-4得分,创下新纪录,并在ROCO数据集上展现出良好泛化能力。
原文摘要 · Abstract (English)
Automating medical reports for retinal images requires a sophisticated blend of visual pattern recognition and deep clinical knowledge. Current Large Vision-Language Models (LVLMs) often struggle in specialized medical fields where data is scarce, leading to models that overfit and miss subtle but critical pathologies. To address this, we introduce DREAM (Dynamic Retinal Enhancement with Adaptive Multi-modal Fusion), a novel framework for high-fidelity medical report generation that excels even with limited data. DREAM employs a unique two-stage fusion mechanism that intelligently integrates visual data with clinical keywords curated by ophthalmologists. First, the Abstractor module maps image and keyword features into a shared space, enhancing visual data with pathology-relevant insights. Next, the Adaptor performs adaptive multi-modal fusion, dynamically weighting the importance of each modality using learnable parameters to create a unified representation. To ensure the model's outputs are semantically grounded in clinical reality, a Contrastive Alignment module aligns these fused representations with ground-truth medical reports during training. By combining medical expertise with an efficient fusion strategy, DREAM sets a new state-of-the-art on the DeepEyeNet benchmark, achieving a BLEU-4 score of 0.241, and further demonstrates strong generalization to the ROCO dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。