arXiv:2511.09893cs.CVcs.CL2025-11

增强关键区域注意力的医学影像描述生成模型,更准且可解释。

Regional Attention-Enhanced Swin Transformer for Clinically Relevant Medical Image Captioning

  • 在Swin-BART中加入轻量级区域注意力模块,聚焦诊断重要区域。
  • ROUGE达0.603,BERTScore达0.807,优于现有方法。
  • 生成结果清晰可解释,适合临床辅助与人类审核场景。

自动化医学图像描述可将复杂放射影像转化为诊断性叙述,辅助报告流程。我们提出一种基于Swin-BART的编码器-解码器系统,引入轻量级区域注意力模块,在跨注意力前放大诊断显著区域。在ROCO数据集上训练评估,模型在保持紧凑与可解释性的同时,达到最先进的语义保真度。结果为三次随机种子的均值±标准差,并包含95%置信区间。相比基线,本方法在ROUGE(提出0.603,ResNet-CNN 0.356,BLIP2-OPT 0.255)和BERTScore(提出0.807,BLIP2-OPT 0.645,ResNet-CNN 0.623)上均有提升,同时在BLEU、CIDEr和METEOR上表现良好。我们还进行了消融实验(开启/关闭区域注意力、词数扫描)、模态分析(CT/MRI/X-ray)、配对显著性检验及定性热图可视化,展示每段描述驱动的关键区域。解码采用束搜索(束宽=4),长度惩罚=1.1,无重复n元组大小=3,最大长度=128。所提设计生成准确、符合临床表达的描述,并提供透明的区域归因,支持安全的研究应用与人工介入。

原文摘要 · Abstract (English)

Automated medical image captioning translates complex radiological images into diagnostic narratives that can support reporting workflows. We present a Swin-BART encoder-decoder system with a lightweight regional attention module that amplifies diagnostically salient regions before cross-attention. Trained and evaluated on ROCO, our model achieves state-of-the-art semantic fidelity while remaining compact and interpretable. We report results as mean$\pm$std over three seeds and include $95\%$ confidence intervals. Compared with baselines, our approach improves ROUGE (proposed 0.603, ResNet-CNN 0.356, BLIP2-OPT 0.255) and BERTScore (proposed 0.807, BLIP2-OPT 0.645, ResNet-CNN 0.623), with competitive BLEU, CIDEr, and METEOR. We further provide ablations (regional attention on/off and token-count sweep), per-modality analysis (CT/MRI/X-ray), paired significance tests, and qualitative heatmaps that visualize the regions driving each description. Decoding uses beam search (beam size $=4$), length penalty $=1.1$, $no\_repeat\_ngram\_size$ $=3$, and max length $=128$. The proposed design yields accurate, clinically phrased captions and transparent regional attributions, supporting safe research use with a human in the loop.

医学影像图像描述注意力机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。