通过细粒度提示引导,提升肺部CT报告中的病灶定位准确性。
Enhancing Fine-Grained Spatial Grounding in 3D CT Report Generation via Discriminative Guidance

- 从报告中提取病灶位置线索,用提示词引导模型生成更精准的空间描述。
- 在CT-RATE上病灶定位F1提升至0.603(相对提高20%),跨数据集表现提升89%。
- 设计分层提问评估协议,揭示现有模型仍难准确识别病灶具体位置。
用于放射科报告生成的视觉-语言模型可从三维扫描生成完整胸部CT报告,有望提升放射科工作流程效率与一致性。然而,现有方法存在两大局限:(i) 训练监督通常较粗粒度,将整个CT体积与完整自由文本报告对齐,缺乏对细粒度属性或病灶位置的显式对齐;(ii) 评估常为整体性指标(如词汇重叠、实体匹配或大模型评判得分),无法诊断空间定位能力。本文提出一种即插即用框架DCP-PD,通过从自由文本报告中提炼细粒度提示,并利用提示丢弃机制缓解模型对捷径依赖。DCP-PD在CT-RATE上将宏平均F1从0.501提升至0.603(相对提升20%),并在Rad-ChestCT上将F1从0.266大幅提升至0.503(相对提升89%)。最后,我们引入分层、位置感知的问题评估协议(存在性→侧别→肺叶),直接检验病灶定位能力,结果显示即使在当前基准上表现优异的模型,其空间定位仍面临挑战。
原文摘要 · Abstract (English)
Vision--language models (VLMs) for radiology report generation (RRG) can produce long-form chest CT reports from volumetric scans and show strong potential to improve radiology workflow efficiency and consistency. However, existing methods face two key limitations: (i) training supervision is often coarse, aligning a whole CT volume with a full free-text report without explicit alignment for fine-grained attributes or pathology locations; and (ii) evaluation is typically holistic (lexical overlap, entity matching, or LLM-as-a-judge scores) and not diagnostic for spatial grounding. We propose \emph{Discriminative Cue-Prompting with Prompt Dropout (DCP-PD)}, a plug-and-play framework that distills fine-grained cues from free-text reports and uses them to guide report generation while mitigating shortcut reliance via prompt dropout. DCP-PD achieves state-of-the-art performance on CT-RATE, improving macro F1 from $=0.501$ to $0.603$ (20% relative), and substantially boosts out-of-distribution performance on Rad-ChestCT from F1 $=0.266$ to $0.503$ (89% relative). Finally, we introduce a hierarchical, location-aware question-set protocol (presence $\rightarrow$ laterality $\rightarrow$ lobe) to directly assess pathology-location grounding, showing that fine-grained spatial localization remains challenging even for models that score highly on current benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。