arXiv:2604.10410cs.AI2026-04中稿 · MIDL 2026被引 2

通过类别对比解码提升胸片报告生成准确性

CWCD: Category-Wise Contrastive Decoding for Structured Medical Report Generation

论文配图:CWCD: Category-Wise Contrastive Decoding for Structured Medical Report Generation
图 1 · 摘自论文原文
  • 按病灶类别分别对比正常与掩蔽图像,生成更精准报告
  • 在多个指标上优于基线方法,减少虚假病灶关联
  • 适合医学影像报告生成、多模态模型优化研究者

胸片解读因解剖结构重叠和病灶表现细微而极具挑战,即使经验丰富的放射科医生也需耗费大量时间。当前聚焦放射领域的基础模型如LLaVA-Rad和Maira-2已将多模态大语言模型(MLLMs)置于自动化胸片报告生成(RRG)的前沿。然而,现有模型采用单次前向传播解码,导致对视觉标记的关注减弱,生成过程愈发依赖语言先验,从而引入虚假病灶共现问题。为此,我们提出类别对比解码(CWCD),一种新颖且模块化的框架,用于增强结构化胸片报告生成(SRRG)。该方法引入类别特定参数化,通过类别特异性视觉提示,对比正常与掩蔽胸片生成类别级报告。实验表明,CWCD在临床有效性和自然语言生成指标上均持续优于基线方法。消融实验证实了各组件对整体性能的贡献。

原文摘要 · Abstract (English)

Interpreting chest X-rays is inherently challenging due to the overlap between anatomical structures and the subtle presentation of many clinically significant pathologies, making accurate diagnosis time-consuming even for experienced radiologists. Recent radiology-focused foundation models, such as LLaVA-Rad and Maira-2, have positioned multi-modal large language models (MLLMs) at the forefront of automated radiology report generation (RRG). However, despite these advances, current foundation models generate reports in a single forward pass. This decoding strategy diminishes attention to visual tokens and increases reliance on language priors as generation proceeds, which in turn introduces spurious pathology co-occurrences in the generated reports. To mitigate these limitations, we propose Category-Wise Contrastive Decoding (CWCD), a novel and modular framework designed to enhance structured radiology report generation (SRRG). Our approach introduces category-specific parameterization and generates category-wise reports by contrasting normal X-rays with masked X-rays using category-specific visual prompts. Experimental results demonstrate that CWCD consistently outperforms baseline methods across both clinical efficacy and natural language generation metrics. An ablation study further elucidates the contribution of each architectural component to overall performance.

医学报告生成多模态模型对比学习胸片分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。