arXiv:2606.01537cs.CVcs.LG2026-06中稿 · ICML

让胸部X光模型学会看心电图和化验数据,提升诊断准确性。

PaCX-MAE: Physiology-Augmented Chest X-Ray Masked Autoencoder

论文配图:PaCX-MAE: Physiology-Augmented Chest X-Ray Masked Autoencoder
图 1 · 摘自论文原文
  • 用生理数据引导X光图像自编码,训练时融合多模态信息。
  • 在多个任务上表现更优,如心电相关任务提升6.5%的F1值。
  • 仅需1%标注数据就有效,适合医疗数据稀缺场景。

临床诊断常需结合影像与生理数据,但现有模型多仅处理单一模态。我们提出PaCX-MAE,一种跨模态蒸馏框架,在训练时注入生理先验,同时推理时仍保持单模态。该方法在域内掩码自编码基础上引入双重对比预测目标,使胸部X光表征与配对的心电图和检验嵌入对齐。在九个基准测试中,其性能持续优于领域特定的MAE,尤其在依赖生理信息的任务上表现突出(如MedMod上AUROC提升2.7;VinDr上F1提升6.5)。该方法在1%标签效率下仍具优势,并在分割任务上与MAE持平,保持解剖结构保真度。零样本及注意力分析表明,模型成功学习关注心脏轮廓等生理指示特征,这些在标准视觉预训练中并不存在。

原文摘要 · Abstract (English)

Clinical diagnosis often requires combining imaging with physiological measurements, yet deployed models typically operate on unimodal data. We present PaCX-MAE, a cross-modal distillation framework that injects physiological priors into chest X-ray (CXR) encoders while remaining strictly unimodal at inference. PaCX-MAE augments in-domain masked autoencoding with a dual contrastive-predictive objective, aligning CXR representations with paired ECG and laboratory embeddings. Extensive evaluation across nine benchmarks demonstrates consistent improvements over domain-specific MAE, particularly on physiology-dependent tasks (e.g., +2.7 AUROC on MedMod; +6.5 F1 on VinDr). The method proves highly label-efficient in the 1% regime and preserves anatomical fidelity, achieving parity with MAE on segmentation tasks. Zero-shot and attention analyses confirm that PaCX-MAE successfully learns to attend to physiological indicators, such as the cardiac silhouette, absent in standard visual pretraining.

医学影像多模态自编码器小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。