融合影像、结构与语义三模态信息,精准预测肺部病变严重程度并给出可信度。
TMF-RSE: Tri-Modal Fusion with Regional Semantics and Evidential Uncertainty for Lung Severity Scoring

- 三模态融合:结合2D影像、分割掩码与视觉语言模型语义特征
- 在两个数据集上实现低于4.02的平均绝对误差和高于0.96的皮尔逊相关性
- 输出预测结果同时提供不确定性估计,适合临床辅助决策场景
从胸部影像中准确量化肺部疾病严重程度对临床决策和资源分配至关重要。我们提出一种三模态深度学习框架TMF-RSE(Tri-Modal Fusion with Regional Semantics and Evidential Uncertainty),融合二维胸片输入的外观特征、肺部分割掩码的结构特征以及视觉语言模型(VLMs)的语义特征,用于严重程度量化。该方法采用互补融合机制,整合语义引导、结构先验及模态间的层次化交互。模型使用证据回归,同时输出严重程度预测值与不确定性估计。在Per-COVID-19 CT和RALO数据集上的实验表明,TMF-RSE优于近期基于Transformer的基线模型,在Per-COVID-19验证集上达到4.02的MAE和0.9629的皮尔逊相关性,在RALO地理范围评估中实现0.339的MAE和0.973的皮尔逊相关性。
原文摘要 · Abstract (English)
Accurate quantification of lung disease severity from chest imaging is critical for clinical decision-making and resource allocation. We propose a tri-modal deep learning framework, TMF-RSE (Tri-Modal Fusion with Regional Semantics and Evidential Uncertainty), that combines appearance features from two-dimensional chest inputs, structural features from lung segmentation masks, and semantic features from vision-language models (VLMs) for severity quantification. Our approach employs complementary fusion mechanisms that integrate semantic guidance, structural priors, and hierarchical interactions across modalities. The model employs evidential regression to provide both severity predictions and uncertainty estimates. Experiments on the Per-COVID-19 CT and RALO datasets show that TMF-RSE outperforms recent transformer-based baselines, achieving MAE of 4.02 and Pearson correlation of 0.9629 on Per-COVID-19 validation, and 0.339 MAE / 0.973 PC on RALO geographic extent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。