arXiv:2506.23102eess.IVcs.CV2025-06中稿 · ECCV被引 10

让3D CT报告生成更精准,聚焦病灶区域并引导模型关注关键解剖结构。

Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation

  • 用区域感知的SlowFast分词器提取医学有意义区域特征。
  • 通过伪掩码引导模型关注诊断关键部位,提升定位准确性。
  • 结合病灶大小、位置等结构化信息生成临床可信报告,适合医学AI研究者。

当前的CT报告生成框架多依赖全局特征表示,难以捕捉区域特异性细节,可能遗漏异常。为此,我们提出MedRegion-CT,一个以区域为中心的多模态大语言模型框架,包含三项创新:首先,重新引入SlowFast策略,设计基于区域的SlowFast分词器,从临床上有意义的区域中提取特征;其次,生成伪掩码引导模型关注具有诊断意义的解剖区域,促进对整体扫描上下文的系统理解;第三,将病灶的定量信息(如大小、直径、空间位置)编码为结构化文本提示,实现上下文感知与临床驱动的报告生成。为实现严格评估,我们在多机构结构化报告生成基准上验证该框架。实验结果表明,MedRegion-CT在语言质量和临床准确性方面均达到当前最优水平。代码已公开于:https://github.com/babbu3682/MedRegion-CT。

原文摘要 · Abstract (English)

Current CT report generation frameworks predominantly rely on global feature representations, often failing to capture region-specific details and potentially missing certain abnormalities. To overcome this limitation, we propose MedRegion-CT, a region-focused multimodal large language model framework featuring three key innovations. First, we revisit the SlowFast strategy to jointly model global and fine-grained information and adapt it to the medical domain via a Region-based SlowFast Tokenizer that extracts tokens guided by clinically meaningful regions. Second, generated pseudo-masks guide the model to attend to diagnostically important anatomical regions, facilitating a systematic understanding of the overall scan context. Third, quantitative lesion information, including size, diameter, and spatial location, is encoded as structured textual prompts, enabling context-aware and clinically informed report generation. To enable rigorous evaluation, we validate our framework on multi-institutional structured report generation benchmarks. Experimental results demonstrate that MedRegion-CT achieves state-of-the-art performance, outperforming existing approaches in both linguistic quality and clinical accuracy. All code is publicly available at: https://github.com/babbu3682/MedRegion-CT.

CT报告生成多模态大模型区域感知医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。