让CT报告生成可预测、可控制,提升医学影像理解能力
SliceWorld: A Predictive and Controllable World-State Model for CT Report Generation

- 将CT切片按纵向序列建模,编码解剖与病灶因子状态
- 支持未来切片预测、病灶因子干预和报告可控生成
- 适合医学AI研发者和临床辅助系统开发者
CT报告生成需整合数百张轴向切片的三维解剖结构与病灶信息。现有方法多采用图像到文本的直接映射,难以建模切片间证据演变或报告对病灶因子变化的响应。我们提出SliceWorld,一种针对CT的全局状态框架,将轴向CT扫描视为沿z轴有序序列。该模型将前序切片证据编码为包含解剖、病灶与不确定性成分的因子感知隐状态,并投影为用于多步未来切片特征预测、病灶因子干预及大语言模型报告生成的世界令牌。模型先在切片序列上通过预测、因子感知与反事实目标预训练,再在配对的CT-报告数据上微调。在M3D-Cap和CT-RATE数据集上的实验表明,SliceWorld在自然语言生成指标与临床导向评估中均表现更优。进一步分析显示其具备多时步未来切片预测能力、可测量的因子对齐、少切片鲁棒性以及选择性病灶敏感的报告调节能力。
原文摘要 · Abstract (English)
CT report generation (CTRG) requires models to summarize three-dimensional anatomical context and pathological findings from hundreds of axial slices. Existing methods typically learn a direct image-to-text mapping, providing limited mechanisms for modeling how CT evidence evolves across slices or how reports respond to controlled changes in latent lesion-related factors. We propose SliceWorld, a CT-specific world-state framework that treats an axial CT scan as an ordered sequence along the z-axis. SliceWorld encodes prefix CT evidence into factor-aware latent states containing anatomy, lesion, and uncertainty components, and projects these states into world tokens used for multi-step future-slice feature prediction, lesion-factor intervention, and LLM-based report generation. The model is first pretrained on CT slice sequences with predictive, factor-aware, and counterfactual objectives, and is then fine-tuned on paired CT-report data. Experiments on M3D-Cap and CT-RATE show that SliceWorld improves natural language generation metrics and clinically oriented automatic evaluation. Further analyses demonstrate multi-horizon future-slice prediction, measurable factor alignment, reduced-slice robustness, and selective lesion-sensitive report modulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。