arXiv:2608.24121cs.CV2026-08

用疾病图谱提升医学影像报告生成的临床准确性

Graph-Supervised Hierarchical Clinical Alignment for Radiology Report Generation with Large Language Models

论文配图:Graph-Supervised Hierarchical Clinical Alignment for Radiology Report Generation with Large Language Models
图 1 · 摘自论文原文
  • 构建疾病驱动的分层对齐机制,分两层细化监督信号
  • 在三个数据集上均优于现有方法,3B模型超越更大模型
  • 适合医疗AI研发者和医学影像报告系统开发者

放射科报告生成(RRG)虽借助大语言模型显著提升流畅性,但临床真实性仍受挑战,因当前监督多在报告整体层面,与报告中基于疾病的细粒度发现存在粒度不匹配。为此,本文提出图监督的分层临床对齐方法,将图像-报告监督重构为疾病条件下的分层对齐任务:疾病中心对齐实现病灶级精准对应,全局临床语义对齐保证报告整体语义连贯。利用临床知识图谱作为训练时结构先验,定义疾病相关监督单元及其临床关系,推理时无额外开销。针对共享病理导致的假负例问题,结合实例条件判别匹配与疾病条件软正则化,实现细粒度且临床一致的跨模态表征。在MIMIC-CXR、IU-Xray和COV-CTR数据集上的实验表明,该方法在常规与临床指标上均持续提升性能,其3B模型超越多个使用7B/13B骨干的大模型,表明优化监督结构比扩大模型规模更有效。

原文摘要 · Abstract (English)

Radiology report generation (RRG) has recently benefited from large language models, which substantially improve report fluency. However, clinically faithful generation remains challenging because current supervision is still imposed mostly at the report level. This creates a granularity mismatch: radiology reports are composed of disease-grounded findings, while existing methods are trained mainly with whole-report objectives. To address this problem, we propose Graph-Supervised Hierarchical Clinical Alignment, which reformulates image-report supervision as a hierarchical clinical alignment problem. Our method structures this alignment as a disease-conditioned process, where supervision is decomposed into two levels: Disease-Centric Alignment for fine-grained disease-specific correspondence, and Global Clinical Semantic Alignment for report-level semantic coherence. A clinical knowledge graph is used as a training-time-only structural prior that defines disease-specific supervision units and their clinical relationships, introducing no additional overhead at inference. Because standard contrastive alignment could produce false negatives when studies share overlapping pathologies, we combine instance-conditioned discriminative matching with disease-conditioned soft regularization, enabling fine-grained yet clinically consistent cross-modal representations. Experiments on MIMIC-CXR, IU-Xray, and COV-CTR show that our method consistently improves performance on both conventional and clinical metrics. Notably, our 3B model surpasses several prior systems with larger 7B/13B backbones, suggesting that improving supervision structure, rather than increasing model size, can be more effective for RRG.

医学报告生成知识图谱多模态对齐临床一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。