用图神经网络增强医学报告生成的拓扑推理能力
Graph-Augmented Topological Internalization with Dual-Stream Classifiers for Medical Report Generation

- 构建图模型捕捉疾病共现关系,实现拓扑知识内化
- 双流分类器提升不平衡数据下的诊断准确率
- 临床语义引导视觉注意力,减少生成幻觉
自动化医学报告生成(MRG)对减轻放射科医生负担、提升诊断效率具有重要意义。现有方法多将不同胸部异常视为孤立分类目标,忽视疾病共现关系,难以显式建模医学拓扑结构,限制了对复杂或细微病灶的推理能力。为此,我们提出图增强的双流医学报告生成框架(GDMRG)。该框架引入拓扑知识内化模块(TKI),利用图卷积网络(GCN)基于全局疾病共现先验生成参数化权重矩阵,实现无需外部检索的高效拓扑知识注入。在此基础上,设计双流分类系统:主分支在拓扑约束下生成诊断提示,辅分支采用非对称优化动态校准高度不平衡样本的决策边界。同时,为建立诊断与视觉定位间的逻辑闭环,设计诊断驱动的空间注意力机制(DGSA),利用高维临床语义重校视觉编码器,缓解特征幻觉。在MIMIC-CXR数据集上的实验表明,GDMRG在保持自然语言流畅性的同时,达到有竞争力的临床效能(CE)。此外,模型在IU X-Ray数据集上展现稳健的零样本泛化能力。本工作提出了一种集成且可解释的医学报告生成范式。
原文摘要 · Abstract (English)
Automated medical report generation, MRG, holds substantial value for alleviating radiologist workload and enhancing diagnostic efficiency. However, mainstream approaches typically treat diverse chest abnormalities as isolated classification targets. This paradigm often overlooks inherent disease co-occurrences and struggles to translate medical topological structures into explicit data correlations, constraining the model's reasoning capacity on complex or subtle lesions. To address this, we propose a Graph-Augmented Dual-Stream Medical Report Generation with Topological Internalization, GDMRG. Our framework introduces a Topological Knowledge Internalization module, TKI, which leverages a Graph Convolutional Network, GCN, to generate an explicit parameterized weight matrix based on global disease co-occurrence priors. This facilitates efficient topological knowledge injection without relying on external retrieval mechanisms. Building upon this, we construct a dual-stream classification system: the main branch generates discrete diagnostic prompts under topological constraints, while the auxiliary branch employs an asymmetric optimization strategy to dynamically calibrate decision boundaries for highly imbalanced samples. Concurrently, to establish a logical closed loop between diagnosis and visual grounding, we design a diagnostic-driven Diagnosis-Guided Spatial Attention, DGSA, that utilizes high-dimensional clinical semantics to recalibrate the visual encoder, mitigating feature hallucinations. Comprehensive experiments on the MIMIC-CXR dataset demonstrate that GDMRG achieves competitive clinical efficacy, CE, while maintaining natural language fluency. Furthermore, our model exhibits robust zero-shot generalization on the IU X-Ray dataset. In summary, this work presents an integrated and interpretable paradigm for medical report generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。