只关注病灶关键区域,比看全图更利于生成准确的放射科报告。
Less Is More? Selective Visual Attention to High-Importance Regions for Multimodal Radiology Summarization
- 聚焦病理相关图像片段,而非整张图像
- 在MIMIC-CXR上实现29.25% BLEU-4和69.83% ROUGE-L
- 适合需要高精度医学影像摘要的研究与临床场景
自动放射科报告摘要旨在将冗长的发现浓缩为简洁的临床印象,但现有多模态模型常受视觉噪声干扰,在发现→印象转换任务中难以超越强文本基线。我们挑战两个普遍假设:(1)更多视觉输入总是更好;(2)当发现已包含丰富图像信息时,多模态模型价值有限。在MIMIC-CXR基准上的控制性消融实验表明,选择性关注病灶相关视觉块比使用全图表现更优。我们提出ViTAS——一种多阶段流水线,结合集成引导的MedSAM2肺部分割、双向交叉注意力多视图融合、基于Shapley值的自适应块聚类及层次化视觉标记化,并输入ViT。ViTAS在多项指标上达到当前最优(BLEU-4: 29.25%,ROUGE-L: 69.83%),定性分析显示事实一致性提升,专家评分最高。结果表明,少而精的视觉输入不仅足够,而且更优。
原文摘要 · Abstract (English)
Automated radiology report summarization aims to distill verbose findings into concise clinical impressions, but existing multimodal models often struggle with visual noise and fail to meaningfully improve over strong text-only baselines in the FINDINGS $\to$ IMPRESSION transformation. We challenge two prevailing assumptions: (1) that more visual input is always better, and (2) that multimodal models add limited value when findings already contain rich image-derived detail. Through controlled ablations on MIMIC-CXR benchmark, we show that selectively focusing on pathology-relevant visual patches rather than full images yields substantially better performance. We introduce ViTAS, Visual-Text Attention Summarizer, a multi-stage pipeline that combines ensemble-guided MedSAM2 lung segmentation, bidirectional cross-attention for multi-view fusion, Shapley-guided adaptive patch clustering, and hierarchical visual tokenization feeding a ViT. ViTAS achieves SOTA results with 29.25% BLEU-4 and 69.83% ROUGE-L, improved factual alignment in qualitative analysis, and the highest expert-rated human evaluation scores. Our findings demonstrate that less but more relevant visual input is not only sufficient but superior for multimodal radiology summarization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。