arXiv:2609.04357eess.IVcs.AI2026-09

多模态AI可自动分级胸片严重程度并生成可解释图像,提升急诊效率。

Cross-modal triage network: a multimodal deep learning framework for severity-based triage and visual explainability in chest radiographs

论文配图:Cross-modal triage network: a multimodal deep learning framework for severity-based triage and visual explainability in chest radiographs
图 1 · 摘自论文原文
  • 融合视觉与文本的跨模态网络,通过门控交叉注意力联合分析影像与报告。
  • 在4级严重度分级上达到0.934的加权肯德尔系数,病理检测宏平均AUC达0.997。
  • 生成的热力图54.3%具备临床可接受的空间定位能力,适合辅助急诊分诊。

目的:胸片数量激增导致分诊瓶颈,急症检查被延迟。现有AI工具多为单模态二分类器,缺乏严重度感知,且多模态系统极少与专家放射科医师进行基准对比。为此,我们开发了一种多模态深度学习框架,实现严重度分诊、病灶检测与原生可视化解释。方法:提出跨模态分诊网络(CMTN),将Swin Transformer V2视觉编码器与PubMedBERT文本编码器通过门控交叉注意力融合。CMTN在34,639对图像-文本数据(12,489名患者)上训练,使用序数焦点损失优化四层级严重度分诊,二元交叉熵损失用于14种病灶检测。除定量基准测试外,还通过两阶段临床审计评估注意力热图:第一阶段对比模型分诊结果与盲法专家严重度评估(100例),第二阶段评估空间-语义一致性(116张热图)。结果:CMTN在严重度分级上与参考标签呈现强序数一致性(二次加权肯德尔系数[QWK] = 0.9341,95% CI: 0.9219–0.9449),14种病灶检测宏平均AUROC达0.9970,延迟仅34~ms,优于当前最优的BioViL多模态基线(QWK = 0.7679)。然而,盲法第一阶段临床审计显示,模型与真实放射科医师判断一致性较低(QWK = 0.1399)。第二阶段发现54.3%的热图达到临床可接受的空间定位水平。结论:CMTN展示了高效多模态架构在胸片分诊中的潜力。算法与放射科医师判断的偏差表明,仅以NLP生成标签作为基准不足以支撑临床部署,强调在实际应用前需依赖放射科医师标注的真实标准。

原文摘要 · Abstract (English)

Purpose: Increased number of chest radiograph (CXR) scans create a triage bottleneck, queueing urgent examinations behind routine ones. Existing AI tools are predominantly unimodal binary classifiers lacking severity awareness, and multimodal systems are rarely benchmarked against expert radiologists. To this end, we developed a multimodal deep learning framework for joint severity triage, pathology detection, and native visual explanation. Approach: We propose the cross-modal triage network (CMTN), fusing a Swin Transformer V2 visual encoder with a PubMedBERT text encoder via gated cross-attention. The CMTN was trained on 34,639 image-text pairs (12,489 patients) from MIMIC-CXR-JPG, optimizing an ordinal focal loss for four-tier severity triage and binary cross-entropy for 14 pathologies. Beyond quantitative benchmarking, attention heatmaps were evaluated against a blinded expert radiologist in a two-phase clinical audit comparing model triage output to expert severity assessment (100 cases) and grading spatial-semantic concordance (116 heatmaps). Results: The CMTN achieved strong ordinal agreement with reference labels (quadratic weighted kappa [QWK] = 0.9341, 95\% CI: 0.9219 to 0.9449) and macro-AUROC of 0.9970 across 14 pathologies, with 34~ms latency, outperforming the state-of-the-art BioViL multimodal baseline (QWK = 0.7679). However, the blinded Phase I clinical audit revealed substantially lower agreement with genuine radiologist judgment (QWK = 0.1399). Phase II found 54.3\% of heatmaps achieved clinically acceptable spatial localization. Conclusions: The CMTN demonstrated an efficient multimodal architecture for CXR triage. The divergence between algorithmic and radiologist agreement demonstrates that benchmark performance against NLP-derived labels is insufficient, highlighting the need for radiologist-labeled ground truth before clinical deployment.

医学影像多模态可解释性分诊

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。