arXiv:2605.17071cs.AI2026-05被引 1

用临床知识图谱引导扩散模型生成更准确的放射科报告

AnchorDiff: Topology-Aware Masked Diffusion with Confidence-based Rewriting for Radiology Report Generation

论文配图:AnchorDiff: Topology-Aware Masked Diffusion with Confidence-based Rewriting for Radiology Report Generation
图 1 · 摘自论文原文
  • 引入临床锚点和拓扑感知掩码,实现双向上下文建模
  • 在MIMIC-CXR和MIMIC-RG4上达到当前最优性能
  • 适合关注医学文本生成与知识融合的研究者

放射科报告生成(RRG)旨在从医学图像自动生成临床准确的文本报告。现有方法主要依赖自回归(AR)语言模型,其因果依赖结构限制生成为单向左到右过程,易产生序列偏差,导致模型倾向于遵循高频报告模板而非充分依据图像证据。本文提出AnchorDiff,首个用于RRG的掩码扩散框架,将基于RadGraph的临床锚点融入扩散语言建模中。通过利用双向上下文和迭代优化,缓解固定顺序自回归解码的局限性。具体地,我们设计了一种拓扑感知训练策略,利用RadGraph衍生的实体层次结构,对临床重要标记实施差异化掩码保护和损失权重分配。此外,还提出了推理阶段重写策略,通过扰动测试检测不稳定的已确定标记,并在去噪过程中选择性修正。在MIMIC-CXR和MIMIC-RG4基准上的大量实验表明,AnchorDiff取得当前最优(SOTA)性能,验证了临床锚定掩码扩散在放射科报告生成中的有效性。

原文摘要 · Abstract (English)

Radiology report generation (RRG) aims to automatically produce clinically accurate textual reports from medical images. Existing methods predominantly rely on autoregressive (AR) language models, whose causal dependency structure restricts generation to a unidirectional left-to-right process. This paradigm can induce sequence bias, where models tend to follow stereotypical token orders and high-frequency report templates rather than fully grounding generation in image-specific evidence. In this paper, we propose AnchorDiff, the first masked-diffusion framework for RRG that integrates knowledge-graph-derived clinical anchors into diffusion language modeling. By leveraging bidirectional context and iterative refinement, AnchorDiff mitigates the limitations of fixed-order autoregressive decoding. Specifically, we introduce a topology-aware training strategy that uses RadGraph-derived entity hierarchies to assign clinically important tokens differentiated masking protection and loss weights. We further design an inference-time rewriting strategy that detects unstable committed tokens through perturbation-based testing and selectively revises them during denoising. Extensive experiments on the MIMIC-CXR and MIMIC-RG4 benchmarks demonstrate that AnchorDiff achieves state-of-the-art (SOTA) performance, showing the effectiveness of clinically anchored masked diffusion for radiology report generation.

放射科报告扩散模型知识融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。