arXiv:2504.19596eess.SPcs.LG2025-04被引 12

构建可处理任意缺失模态的生理信号通用模型

Towards Robust Multimodal Physiological Foundation Models: Handling Arbitrary Missing Modalities

  • 通过解耦多模态信号,分离共性与特性特征
  • 在4个下游任务中均达顶尖性能且对缺失模态鲁棒
  • 适合医疗与脑机接口中不完整数据场景

多模态生理信号(如EEG、ECG、EOG、EMG)在医疗和脑机接口中至关重要。现有方法依赖特定架构与数据集定制融合策略,难以学习跨数据集通用表示,也无法在推理时处理缺失模态。为此,我们提出PhysioOmni,一种多模态生理信号分析的基础模型,通过建模同质与异质特征,实现多模态信号解耦与通用表征提取,并兼容任意缺失模态。PhysioOmni采用解耦的多模态分词器,通过模态无关与模态特异性目标进行掩码信号预训练。为适应多样且不完整的模态组合,预训练编码器在下游数据集上通过原型对齐进行弹性微调。在情绪识别、睡眠分期分类、运动预测和心理负荷检测4个下游任务上的大量实验表明,PhysioOmni不仅达到当前最优性能,还对缺失模态保持强鲁棒性。代码与模型权重将公开。

原文摘要 · Abstract (English)

Multimodal physiological signals, such as EEG, ECG, EOG, and EMG, are crucial for healthcare and brain-computer interfaces. While existing methods rely on specialized architectures and dataset-specific fusion strategies, they struggle to learn universal representations that generalize across datasets and handle missing modalities at inference time. To address these issues, we propose PhysioOmni, a foundation model for multimodal physiological signal analysis that models both homogeneous and heterogeneous features to decouple multimodal signals and extract generic representations while maintaining compatibility with arbitrary missing modalities. PhysioOmni trains a decoupled multimodal tokenizer, enabling masked signal pre-training via modality-invariant and modality-specific objectives. To ensure adaptability to diverse and incomplete modality combinations, the pre-trained encoders undergo resilient fine-tuning with prototype alignment on downstream datasets. Extensive experiments on four downstream tasks, emotion recognition, sleep stage classification, motor prediction, and mental workload detection, demonstrate that PhysioOmni achieves state-of-the-art performance while maintaining strong robustness to missing modalities. Our code and model weights will be released.

生理信号多模态基础模型缺失数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。