arXiv:2411.12126cs.LG2024-11被引 13

解决物联网多模态数据分散不完整问题,让模型在残缺数据上仍能高效学习

MMBind: Unleashing the Potential of Distributed and Heterogeneous Data for Multimodal Learning in IoT

  • 用共享模态拼接异构数据,生成伪配对训练集
  • 在数据缺失和分布差异下,性能优于现有方法
  • 适合真实物联网场景中的多模态模型训练

多模态传感系统在现实应用中日益普及。现有方法大多依赖大量同步、完整的多模态数据训练,但在真实物联网场景中,数据由异构节点分散采集,常缺标签且模态不全。本文提出MMBind,一种面向分布式异构物联网数据的多模态学习数据绑定方法。核心思想是通过充分描述性的共享模态,将来自不同源、不完整的模态数据绑定成伪配对数据集用于训练。同时提出加权对比学习处理数据域偏移,并设计自适应多模态架构,支持多种模态组合训练。在十个真实多模态数据集上的评估表明,MMBind在不同数据缺失程度和域偏移条件下均优于现有先进方法,为物联网多模态基础模型训练提供了新可能。

原文摘要 · Abstract (English)

Multimodal sensing systems are increasingly prevalent in various real-world applications. Most existing multimodal learning approaches heavily rely on training with a large amount of synchronized, complete multimodal data. However, such a setting is impractical in real-world IoT sensing applications where data is typically collected by distributed nodes with heterogeneous data modalities, and is also rarely labeled. In this paper, we propose MMBind, a new data binding approach for multimodal learning on distributed and heterogeneous IoT data. The key idea of MMBind is to construct a pseudo-paired multimodal dataset for model training by binding data from disparate sources and incomplete modalities through a sufficiently descriptive shared modality. We also propose a weighted contrastive learning approach to handle domain shifts among disparate data, coupled with an adaptive multimodal learning architecture capable of training models with heterogeneous modality combinations. Evaluations on ten real-world multimodal datasets highlight that MMBind outperforms state-of-the-art baselines under varying degrees of data incompleteness and domain shift, and holds promise for advancing multimodal foundation model training in IoT applications\footnote (The source code is available via https://github.com/nesl/multimodal-bind).

多模态学习物联网数据绑定异构数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。