解决医学影像报告生成三大难题,提升准确性和可解释性。
Medical Report Generation: A Hierarchical Task Structure-Based Cross-Modal Causal Intervention Framework
- 分层任务设计:低、中、高三级任务协同优化
- 在MIMIC-CXR和Chexpert数据集上显著超越现有方法
- 适合医疗AI研究者与临床辅助系统开发者
医学报告生成(MRG)是现代医疗诊断的关键环节,能自动从影像生成报告以减轻放射科医生负担。然而,可靠的病变描述模型面临三大挑战:领域知识理解不足、文本-视觉实体嵌入对齐差,以及跨模态偏见引发的虚假相关。现有工作仅解决单一问题,本文提出新型分层任务分解框架HTSC-CIF,将挑战划分为低、中、高三级任务:1)低级任务:对齐医学实体特征与空间位置,增强视觉编码器的领域知识;2)中级任务:采用前缀语言建模(文本)与掩码图像建模(图像),通过相互引导提升跨模态对齐;3)高级任务:引入跨模态因果干预模块(前门干预),减少混杂因素并提升可解释性。大量实验验证了该框架的有效性,在MIMIC-CXR和Chexpert数据集上显著优于当前最优方法。代码将在论文接收后公开。
原文摘要 · Abstract (English)
Medical Report Generation (MRG) is a key part of modern medical diagnostics, as it automatically generates reports from radiological images to reduce radiologists' burden. However, reliable MRG models for lesion description face three main challenges: insufficient domain knowledge understanding, poor text-visual entity embedding alignment, and spurious correlations from cross-modal biases. Previous work only addresses single challenges, while this paper tackles all three via a novel hierarchical task decomposition approach, proposing the HTSC-CIF framework. HTSC-CIF classifies the three challenges into low-, mid-, and high-level tasks: 1) Low-level: align medical entity features with spatial locations to enhance domain knowledge for visual encoders; 2) Mid-level: use Prefix Language Modeling (text) and Masked Image Modeling (images) to boost cross-modal alignment via mutual guidance; 3) High-level: a cross-modal causal intervention module (via front-door intervention) to reduce confounders and improve interpretability. Extensive experiments confirm HTSC-CIF's effectiveness, significantly outperforming state-of-the-art (SOTA) MRG methods. Code will be made public upon paper acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。