arXiv:2504.09707cs.AIcs.IT2025-04被引 8

用少量数据对实现多模态信号高效对齐,提升物联网感知性能

InfoMAE: Pair-Efficient Cross-Modal Alignment for Multimodal Time-Series Sensing Signals

  • 基于信息论设计新框架,同时解决分布与实例级对齐
  • 仅用少量数据对,下游任务性能提升超60%
  • 适合数据稀疏的物联网多模态场景,尤其单模态数据多

标准多模态自监督学习算法将跨模态同步视为预训练中的隐式标签,对多模态样本的数量和质量要求高。这在物联网应用中严重制约了感知智能表现,因时间序列信号异构性与不可解释性导致单模态数据丰富但高质量多模态配对稀缺。本文提出InfoMAE,一种跨模态对齐框架,通过促进预训练单模态表示的高效跨模态对齐,在自监督设定下解决多模态配对效率问题。InfoMAE采用新颖的信息论启发公式,在有限数据对下实现高效的跨模态对齐。在两个真实世界物联网应用上的大量实验表明,该方法显著提升了多模态配对效率,使下游多模态任务性能提升超过60%,同时平均提升单模态任务准确率22%。

原文摘要 · Abstract (English)

Standard multimodal self-supervised learning (SSL) algorithms regard cross-modal synchronization as implicit supervisory labels during pretraining, thus posing high requirements on the scale and quality of multimodal samples. These constraints significantly limit the performance of sensing intelligence in IoT applications, as the heterogeneity and the non-interpretability of time-series signals result in abundant unimodal data but scarce high-quality multimodal pairs. This paper proposes InfoMAE, a cross-modal alignment framework that tackles the challenge of multimodal pair efficiency under the SSL setting by facilitating efficient cross-modal alignment of pretrained unimodal representations. InfoMAE achieves \textit{efficient cross-modal alignment} with \textit{limited data pairs} through a novel information theory-inspired formulation that simultaneously addresses distribution-level and instance-level alignment. Extensive experiments on two real-world IoT applications are performed to evaluate InfoMAE's pairing efficiency to bridge pretrained unimodal models into a cohesive joint multimodal model. InfoMAE enhances downstream multimodal tasks by over 60% with significantly improved multimodal pairing efficiency. It also improves unimodal task accuracy by an average of 22%.

多模态学习自监督物联网时间序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。