无需标注,通过解剖结构自监督学习生成可解释的医学报告
Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation
- 用文本提示构建解剖层级图,递归重建解剖区域以对齐视觉与文本
- 在无专家标注下,词汇准确率提升10%,临床有效性提高25%
- 适合医疗AI、医学影像分析和可解释报告生成的研究者
视觉引导的医学报告生成旨在基于显式视觉证据生成临床准确的描述,提升可解释性并便于融入临床流程。然而现有方法通常依赖需大量专家标注的检测模块,导致标注成本高,且因病理分布偏差限制泛化能力。为此,本文提出自监督解剖一致性学习(SS-ACL)——一种无需标注的框架,利用简单文本提示对齐生成报告与对应解剖区域。该方法基于人体解剖学固有的上下文包含结构,构建分层解剖图谱,按空间位置组织实体,并递归重建细粒度解剖区域,以强制样本内空间对齐,引导注意力聚焦于文本提示的视觉相关区域。为进一步增强跨样本异常识别的语义对齐,引入基于解剖一致性的区域级对比学习。对齐嵌入作为报告生成先验,使注意力图提供可解释的视觉证据。大量实验表明,无需专家标注的SS-ACL在报告生成上优于当前最优方法:词汇准确率提升10%,临床有效性提升25%;同时在下游视觉任务中表现优异,零样本视觉定位性能超越当前领先视觉基础模型8%。
原文摘要 · Abstract (English)
Vision-grounded medical report generation aims to produce clinically accurate descriptions of medical images, anchored in explicit visual evidence to improve interpretability and facilitate integration into clinical workflows. However, existing methods often rely on separately trained detection modules that require extensive expert annotations, introducing high labeling costs and limiting generalizability due to pathology distribution bias across datasets. To address these challenges, we propose Self-Supervised Anatomical Consistency Learning (SS-ACL) -- a novel and annotation-free framework that aligns generated reports with corresponding anatomical regions using simple textual prompts. SS-ACL constructs a hierarchical anatomical graph inspired by the invariant top-down inclusion structure of human anatomy, organizing entities by spatial location. It recursively reconstructs fine-grained anatomical regions to enforce intra-sample spatial alignment, inherently guiding attention maps toward visually relevant areas prompted by text. To further enhance inter-sample semantic alignment for abnormality recognition, SS-ACL introduces a region-level contrastive learning based on anatomical consistency. These aligned embeddings serve as priors for report generation, enabling attention maps to provide interpretable visual evidence. Extensive experiments demonstrate that SS-ACL, without relying on expert annotations, (i) generates accurate and visually grounded reports -- outperforming state-of-the-art methods by 10\% in lexical accuracy and 25\% in clinical efficacy, and (ii) achieves competitive performance on various downstream visual tasks, surpassing current leading visual foundation models by 8\% in zero-shot visual grounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。