融合影像与临床数据,提升胸部X光诊断准确率
Clinically-aligned Multi-modal Chest X-ray Classification
- 用多模态Transformer融合不同视角的X光片和临床信息
- 在MIMIC-CXR和CXR-LT数据集上均超越现有最佳水平
- 适合医疗AI研究者及放射科辅助系统开发者
放射科是现代医疗的核心,但日益增长的需求与人员短缺带来严峻挑战。人工智能有望助力放射科医生应对这些压力。胸部X光分类因广泛应用和临床重要性,非常适合融入医生工作流程。然而,现有方法大多仅依赖单视角图像,忽略了报告时可用的结构化临床信息和多图例数据。本文提出CaMCheX,一种基于多模态Transformer的框架,将多视角胸部X光片与结构化临床数据对齐,更贴近临床诊断逻辑。模型采用视图专用的ConvNeXt编码器处理正位与侧位胸片,再通过Transformer融合模块结合临床指征、病史和生命体征。该设计生成具有上下文感知能力的表征,反映真实临床推理过程。实验结果在原始MIMIC-CXR数据集和最新CXR-LT基准上均超越当前最优水平,证明了临床对齐多模态融合在胸部X光分类中的价值。
原文摘要 · Abstract (English)
Radiology is essential to modern healthcare, yet rising demand and staffing shortages continue to pose major challenges. Recent advances in artificial intelligence have the potential to support radiologists and help address these challenges. Given its widespread use and clinical importance, chest X-ray classification is well suited to augment radiologists' workflows. However, most existing approaches rely solely on single-view, image-level inputs, ignoring the structured clinical information and multi-image studies available at the time of reporting. In this work, we introduce CaMCheX, a multimodal transformer-based framework that aligns multi-view chest X-ray studies with structured clinical data to better reflect how clinicians make diagnostic decisions. Our architecture employs view-specific ConvNeXt encoders for frontal and lateral chest radiographs, whose features are fused with clinical indications, history, and vital signs using a transformer fusion module. This design enables the model to generate context-aware representations that mirror reasoning in clinical practice. Our results exceed the state of the art for both the original MIMIC-CXR dataset and the more recent CXR-LT benchmarks, highlighting the value of clinically grounded multimodal alignment for advancing chest X-ray classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。