arXiv:2602.22917cs.CV2026-02中稿 · CVPR被引 2

少标签下提升多模态模型跨域泛化能力

Towards Multimodal Domain Generalization with Few Labels

  • 通过一致性和分歧感知机制,利用少量标签和大量无标签数据训练
  • 在标准和缺模态场景下均显著优于现有方法
  • 适合低资源多模态跨域应用,如医疗影像分析

多模态模型应能在未见领域中保持鲁棒性,同时数据高效以降低标注成本。为此,我们提出并研究了一个新问题:半监督多模态域泛化(SSMDG),旨在用少量标注样本从多源数据中学习鲁棒的多模态模型。现有方法存在三大局限:多模态域泛化无法利用无标签数据,半监督多模态学习忽略域偏移,而半监督域泛化仅限单模态输入。为此,我们提出统一框架,包含三个核心组件:基于可信融合单模态共识的共识驱动一致性正则化,用于生成可靠伪标签;分歧感知正则化,有效利用不一致的非共识样本;跨模态原型对齐,实现域和模态不变表示,并通过跨模态翻译增强缺失模态下的鲁棒性。我们还建立了首个SSMDG基准,实验表明该方法在标准和缺模态场景下均持续优于强基线。代码与数据集已开源。

原文摘要 · Abstract (English)

Multimodal models ideally should generalize to unseen domains while remaining data-efficient to reduce annotation costs. To this end, we introduce and study a new problem, Semi-Supervised Multimodal Domain Generalization (SSMDG), which aims to learn robust multimodal models from multi-source data with few labeled samples. We observe that existing approaches fail to address this setting effectively: multimodal domain generalization methods cannot exploit unlabeled data, semi-supervised multimodal learning methods ignore domain shifts, and semi-supervised domain generalization methods are confined to single-modality inputs. To overcome these limitations, we propose a unified framework featuring three key components: Consensus-Driven Consistency Regularization, which obtains reliable pseudo-labels through confident fused-unimodal consensus; Disagreement-Aware Regularization, which effectively utilizes ambiguous non-consensus samples; and Cross-Modal Prototype Alignment, which enforces domain- and modality-invariant representations while promoting robustness under missing modalities via cross-modal translation. We further establish the first SSMDG benchmarks, on which our method consistently outperforms strong baselines in both standard and missing-modality scenarios. Our benchmarks and code are available at https://github.com/lihongzhao99/SSMDG.

多模态域泛化半监督少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。