arXiv:2603.17718cs.CV2026-03被引 1

用视觉差异提示提升大模型生成胸部CT报告的准确性

DiffVP: Differential Visual Semantic Prompting for LLM-Based CT Report Generation

  • 通过提取扫描与参考图像的语义差异,生成可学习的视觉提示
  • 在两个基准上提升BLEU-1-4平均得分10.98和4.36,F1达0.421
  • 适合需要精准报告生成且关注病灶对比的临床研究者

尽管大语言模型已推动CT报告生成发展,但现有方法通常整体编码三维体积,难以区分信息性线索与冗余解剖背景。受放射科认知减法启发,本文提出差分视觉提示(DiffVP),让报告生成基于显式的高层语义扫描-参考差异,而非仅依赖绝对视觉特征。DiffVP采用分层差异提取器,将互补的全局与局部语义差异映射至共享潜在空间,并通过差异到提示生成器将其转换为可学习的视觉前缀令牌以条件化大模型。这些差异提示作为结构化条件信号,隐式抑制不变解剖结构,同时增强诊断相关视觉证据,从而实现无需显式病灶定位的精准报告生成。在两个大规模基准测试中,DiffVP持续优于先前方法,平均BLEU-1-4分别提升10.98和4.36,并在RadGenome-ChestCT上进一步提升临床有效性(F1分数0.421)。所有代码将公开于https://github.com/ArielTYH/DiffVP/。

原文摘要 · Abstract (English)

While large language models (LLMs) have advanced CT report generation, existing methods typically encode 3D volumes holistically, failing to distinguish informative cues from redundant anatomical background. Inspired by radiological cognitive subtraction, we propose Differential Visual Prompting (DiffVP), which conditions report generation on explicit, high-level semantic scan-to-reference differences rather than solely on absolute visual features. DiffVP employs a hierarchical difference extractor to capture complementary global and local semantic discrepancies into a shared latent space, along with a difference-to-prompt generator that transforms these signals into learnable visual prefix tokens for LLM conditioning. These difference prompts serve as structured conditioning signals that implicitly suppress invariant anatomy while amplifying diagnostically relevant visual evidence, thereby facilitating accurate report generation without explicit lesion localization. On two large-scale benchmarks, DiffVP consistently outperforms prior methods, improving the average BLEU-1-4 by +10.98 and +4.36, respectively, and further boosts clinical efficacy on RadGenome-ChestCT (F1 score 0.421). All codes will be released at https://github.com/ArielTYH/DiffVP/.

CT报告生成视觉提示大模型医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。