arXiv:2511.11066cs.CVcs.AI2025-11被引 3

通过分层辅助信号提升影像报告生成的解剖定位精度。

S2D-ALIGN: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report Generation

  • 采用由粗到细的多阶段辅助学习策略,逐步引入不同粒度的引导信号。
  • 在MIMIC-CXR和IU X-Ray数据集上达到当前最优性能,显著提升报告准确性。
  • 适合需要高解剖精准度的医学影像生成任务,如临床辅助诊断系统。

放射科报告生成(RRG)旨在从医学影像自动生成诊断报告。现有方法主要依赖多模态大语言模型(MLLMs)的跨模态生成能力,通过监督微调(SFT)优化图像与报告间的对齐。然而,传统SFT仅在实例级别进行对齐,难以建立解剖学意义上的精确对齐,导致模板化报告影响生成质量。为此,我们提出S2D-ALIGN,一种新型SFT范式,通过引入不同粒度的辅助信号实现解剖学对齐。该方法采用由浅入深的策略:先进行粗粒度图像-报告配对,再引入参考报告进行实例级指导,最终利用关键短语锚定具体解剖结构。为衔接各对齐阶段,设计基于记忆的适配器以实现特征共享。在公开数据集MIMIC-CXR和IU X-Ray上的实验表明,S2D-ALIGN性能优于现有方法。消融实验验证了多阶段、辅助引导策略的有效性,为复杂多模态生成任务中的对齐能力提升提供了新方向。

原文摘要 · Abstract (English)

Radiology Report Generation (RRG) aims to automatically generate diagnostic reports from radiology images. To achieve this, existing methods have leveraged the powerful cross-modal generation capabilities of Multimodal Large Language Models (MLLMs), primarily focusing on optimizing cross-modal alignment between radiographs and reports through Supervised Fine-Tuning (SFT). However, by only performing instance-level alignment with the image-text pairs, the standard SFT paradigm fails to establish anatomically-grounded alignment, where the templated nature of reports often leads to sub-optimal generation quality. To address this, we propose \textsc{S2D-Align}, a novel SFT paradigm that establishes anatomically-grounded alignment by leveraging auxiliary signals of varying granularities. \textsc{S2D-Align} implements a shallow-to-deep strategy, progressively enriching the alignment process: it begins with the coarse radiograph-report pairing, then introduces reference reports for instance-level guidance, and ultimately utilizes key phrases to ground the generation in specific anatomical details. To bridge the different alignment stages, we introduce a memory-based adapter that empowers feature sharing, thereby integrating coarse and fine-grained guidance. For evaluation, we conduct experiments on the public \textsc{MIMIC-CXR} and \textsc{IU X-Ray} benchmarks, where \textsc{S2D-Align} achieves state-of-the-art performance compared to existing methods. Ablation studies validate the effectiveness of our multi-stage, auxiliary-guided approach, highlighting a promising direction for enhancing grounding capabilities in complex, multi-modal generation tasks.

报告生成医学影像多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。