用可解释的概念推理框架,让AI更懂脑瘤治疗评估标准
TRACE: A Concept Bottleneck Model for Longitudinal 3D Glioblastoma Response Assessment

- 基于RANO标准构建概念瓶颈,分步推理肿瘤变化
- 4分类宏F1达0.4769,进展判断宏F1达0.7085
- 医生可干预概念,提升预测可信度,适合临床辅助
纵向胶质母细胞瘤疗效评估需依据RANO标准对比多时点MRI中的细微肿瘤变化。现有深度学习方法直接从影像特征预测结果,缺乏可解释性。本文提出TRACE模型,一种符合RANO 2.0标准的3D概念瓶颈模型,用于纵向3D MRI上的4类胶质母细胞瘤响应分类。该模型使用共享3D视觉编码器处理基线与随访多模态MRI,预测临床有意义的肿瘤测量作为核心概念,通过确定性规则生成下游RANO相关概念,并将扫描间隔与新病灶信息作为旁路概念输入。整个流程将评估转化为结构化概念推理而非图像到标签的直接映射。在LUMIERE数据集上采用5折患者级交叉验证,模型取得4分类宏F1 0.4769,以及进展-非进展二分类宏F1 0.7085,优于基准概念瓶颈模型,且与已有非可解释模型相当。消融实验表明专家定义的RANO图谱和干预一致性训练对性能至关重要;干预实验显示修正概念可提升下游预测效果。结果表明,结构化概念瓶颈为纵向胶质母细胞瘤评估提供了透明且临床对齐的方向,但亟需更大规模协议对齐的数据集与外部验证。
原文摘要 · Abstract (English)
Longitudinal glioblastoma response assessment requires comparing subtle tumor changes across MRI time points using structured clinical criteria such as RANO. However, most deep learning methods predict response labels directly from imaging features, which limits clinical inspection, verification, and correction. We introduce TRACE, a RANO 2.0-aligned concept bottleneck model for interpretable 4-class glioblastoma response classification on longitudinal 3D MRI. TRACE processes paired baseline and follow-up multimodal MRI scans with a shared 3D vision encoder, predicts clinically meaningful tumor measurements as root concepts, computes downstream RANO-derived concepts through deterministic rules, and incorporates scan interval and new-lesion information as passthrough concepts. This design frames response assessment as structured concept reasoning rather than direct image-to-label prediction. Using 5-fold patient-wise cross-validation on the LUMIERE dataset, TRACE achieves a 4-class macro F1 of 0.4769 and a binary progression-versus-non-progression macro F1 of 0.7085. It improves over a concept bottleneck baseline and remains within the range of published non-interpretable deep learning approaches. Ablation studies show that the expert RANO graph and intervention-consistency training are important for performance, while intervention experiments demonstrate that correcting concepts can improve downstream predictions. These results suggest that structured concept bottlenecks offer a transparent and clinically aligned direction for longitudinal glioblastoma response assessment, while highlighting the need for larger protocol-aligned datasets and external validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。