arXiv:2605.30984cs.CVcs.AI2026-05被引 1

解决3D CT报告生成中模板坍缩问题,提升罕见病检出率。

Generating Reports or Repeating Templates? Measuring and Mitigating Template Collapse in 3D CT Report Generation

论文配图:Generating Reports or Repeating Templates? Measuring and Mitigating Template Collapse in 3D CT Report Generation
图 1 · 摘自论文原文
  • 分离诊断与表述,用查询变换器检测病灶,检索临床实例
  • 相比基线,罕见病检出率提升30%以上,宏观F1达0.487
  • 适合医学AI研发者、放射科医生关注临床真实性

现代3D医学视觉语言模型虽能生成流畅的放射科报告,但病理检出率极低且输出高度重复,陷入通用模板的‘模板坍缩’。这一现象源于3D医学影像数据少、标签严重不均衡及体积分解器信号弱等约束,导致文本生成目标诱导捷径学习,产生看似流畅却缺乏临床依据的报告。本文通过临床保真度、输出多样性、正常模板偏差和罕见发现存活率四项指标系统诊断该问题。提出CLarGen框架,将‘说什么’(临床检测)与‘怎么讲’(语言合成)解耦:使用潜在查询变压器进行多标签病灶检测,病理引导检索临床匹配样例,并由医学语言模型综合检测结果与检索上下文生成报告。在多个3D CT报告生成基准上,CLarGen显著缓解模板坍缩,临床准确率大幅提升(宏平均F1从0.189增至0.487,临床报告生成分数从0.368升至0.472),同时保持报告流畅性。结果表明,显式的临床对齐是抵抗模板坍缩的关键。代码将在接受后发布。

原文摘要 · Abstract (English)

Modern 3D medical vision-language models (VLMs) can generate fluent radiology-style text while exhibit critically low pathology detection and output diversity, collapsing to generic templates that under-report rare yet critical findings. We identify this failure mode as Template Collapse. This failure stems from the unique constraints of 3D medical imaging, e.g., limited data, severe label imbalance, and weak signals from volumetric encoders. Under these constraints, text-generation objectives encourage shortcut learning and fluent but weakly grounded reports. We systematically diagnose the Template Collapse through clinical fidelity, output diversity, normal-template bias, and rare-finding survival. To mitigate it, we propose CLarGen, a decoupled framework that separates what to say (clinical detection) from how to say it (language synthesis). CLarGen uses (i) a Latent Query Transformer for multi-label pathology detection, (ii) pathology-guided retrieval for clinically matched exemplars, and (iii) a medical language model to synthesize the final report from detected findings and retrieved context. Across state-of-the-art 3D CT report generation baselines, CLarGen mitigates Template Collapse and substantially improves clinical accuracy (macro-F1 0.487 vs. 0.189; CRG 0.472 vs. 0.368) while maintaining fluent reporting. Our results suggest that explicit, measurable clinical grounding is essential for template-collapse-resistant 3D CT report generation. Code will be released upon acceptance.

3D医学影像报告生成模板坍缩临床对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。