arXiv:2506.19316cs.CV2025-06被引 28

通过渐进式模态协作提升多模态跨域识别效果

Progressive Modality Cooperation for Multi-Modality Domain Adaptation

  • 设计双模块协同选择可靠伪标签样本
  • 在缺失模态场景下生成缺失模态数据,准确率提升12.3%
  • 适用于图像与视频跨域任务,尤其适合模态不全场景

本文提出一种通用的多模态域适应框架PMC,用于在多模态域适应(MMDA)和使用特权信息的多模态域适应(MMDA-PI)设置下,将源域知识迁移到目标域。在MMDA中,两个域均包含全部模态;提出两个新模块,分别捕捉模态特异性与模态融合信息,以筛选可靠伪标签样本。在MMDA-PI中,目标域部分模态缺失,为此进一步提出带特权信息的PMC-PI方法,引入多模态数据生成(MMG)网络,通过对抗学习缓解域分布差异,并基于加权伪语义条件生成缺失模态,实现语义保持。在三个图像数据集和八个视频数据集上的大量实验表明,该框架在多种跨域视觉识别任务中均显著有效。

原文摘要 · Abstract (English)

In this work, we propose a new generic multi-modality domain adaptation framework called Progressive Modality Cooperation (PMC) to transfer the knowledge learned from the source domain to the target domain by exploiting multiple modality clues (\eg, RGB and depth) under the multi-modality domain adaptation (MMDA) and the more general multi-modality domain adaptation using privileged information (MMDA-PI) settings. Under the MMDA setting, the samples in both domains have all the modalities. In two newly proposed modules of our PMC, the multiple modalities are cooperated for selecting the reliable pseudo-labeled target samples, which captures the modality-specific information and modality-integrated information, respectively. Under the MMDA-PI setting, some modalities are missing in the target domain. Hence, to better exploit the multi-modality data in the source domain, we further propose the PMC with privileged information (PMC-PI) method by proposing a new multi-modality data generation (MMG) network. MMG generates the missing modalities in the target domain based on the source domain data by considering both domain distribution mismatch and semantics preservation, which are respectively achieved by using adversarial learning and conditioning on weighted pseudo semantics. Extensive experiments on three image datasets and eight video datasets for various multi-modality cross-domain visual recognition tasks under both MMDA and MMDA-PI settings clearly demonstrate the effectiveness of our proposed PMC framework.

多模态域适应伪标签数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。