用跨模态重建提升生理信号多模态模型表现
Promoting cross-modal representations to improve multimodal foundation models for physiological signals
- 采用掩码自编码预训练,通过跨模态重建增强信息融合
- 输入模态丢弃策略使下游任务性能普遍提升
- 注意力更跨模态且时间对齐,适合医疗多模态研究
许多医疗应用本质上是多模态的,涉及多种生理信号。随着传感器普及,改进多模态医疗数据的机器学习方法至关重要。预训练基础模型是一条有前景的路径,但当前针对医疗领域的基础模型构建仍处于探索阶段,尚不清楚何种预训练策略最有效。这部分归因于多模态健康数据的挑战:跨患者数据获取困难且成本高,个体间差异大,不同模态在下游任务中信息量不均。本文以PhysioNet 2018数据集为实验对象,使用掩码自编码目标预训练多模态模型。结果显示,模型学习到的表示可通过线性探测应用于多种下游任务。我们假设跨模态重建目标对成功多模态训练至关重要,因其促进模态间信息整合。实验证明,输入空间的模态丢弃能提升多个下游任务的表现。同时发现,采用对比学习目标的晚期融合模型在多任务上表现较差。最后分析显示,经预训练后,注意力权重更跨模态且时间对齐,每个神经单元编码的模态分布也更广泛。总体而言,本工作验证了在多样化生理信号源上构建多模态基础模型的有效性,并主张显式引入跨模态机制可增强预训练策略。
原文摘要 · Abstract (English)
Many healthcare applications are inherently multimodal, involving several physiological signals. As sensors for these signals become more common, improving machine learning methods for multimodal healthcare data is crucial. Pretraining foundation models is a promising avenue for success. However, methods for developing foundation models in healthcare are still in early exploration and it is unclear which pretraining strategies are most effective given the diversity of physiological signals. This is partly due to challenges in multimodal health data: obtaining data across many patients is difficult and costly, there is a lot of inter-subject variability, and modalities are often heterogeneously informative across downstream tasks. Here, we explore these challenges in the PhysioNet 2018 dataset. We use a masked autoencoding objective to pretrain a multimodal model. We show that the model learns representations that can be linearly probed for a diverse set of downstream tasks. We hypothesize that cross-modal reconstruction objectives are important for successful multimodal training, as they encourage the model to integrate information across modalities. We demonstrate that modality dropout in the input space improves performance across downstream tasks. We also find that late-fusion models pretrained with contrastive learning objectives are less effective across multiple tasks. Finally, we analyze the model's representations, showing that attention weights become more cross-modal and temporally aligned with our pretraining strategy. The learned embeddings also become more distributed in terms of the modalities encoded by each unit. Overall, our work demonstrates the utility of multimodal foundation models with health data, even across diverse physiological data sources. We further argue that explicit methods for inducing cross-modality may enhance multimodal pretraining strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。