少样本+分布漂移下,多模态协同训练提升模型泛化能力。
Efficient Generalization via Multimodal Co-Training under Data Scarcity and Distribution Shift
- 利用多模态数据与未标注样本,通过模态间一致性促进学习。
- 理论证明迭代协同训练能有效降低分类误差并提升泛化性能。
- 首次量化多模态协同训练中各优势来源,适合数据稀缺场景应用。
本文研究了一种在标签数据有限且存在分布偏移情况下的多模态协同训练框架,旨在提升模型泛化能力。我们深入探讨了该框架的理论基础,推导出利用未标注数据及促进不同模态分类器间一致性的条件,从而显著改善泛化表现。此外,我们进行了收敛性分析,证实迭代协同训练可有效减少分类错误。更重要的是,我们首次在多模态协同训练背景下提出了一个新颖的泛化界,将优势分解并量化为三部分:利用未标注多模态数据、促进跨视图一致性以及维持条件视图独立性。研究结果表明,该方法是一种结构化的数据高效且鲁棒的AI系统构建策略,适用于动态真实环境。理论分析与已有协同训练原则展开对话,并提前建立基础。
原文摘要 · Abstract (English)
This paper explores a multimodal co-training framework designed to enhance model generalization in situations where labeled data is limited and distribution shifts occur. We thoroughly examine the theoretical foundations of this framework, deriving conditions under which the use of unlabeled data and the promotion of agreement between classifiers for different modalities lead to significant improvements in generalization. We also present a convergence analysis that confirms the effectiveness of iterative co-training in reducing classification errors. In addition, we establish a novel generalization bound that, for the first time in a multimodal co-training context, decomposes and quantifies the distinct advantages gained from leveraging unlabeled multimodal data, promoting inter-view agreement, and maintaining conditional view independence. Our findings highlight the practical benefits of multimodal co-training as a structured approach to developing data-efficient and robust AI systems that can effectively generalize in dynamic, real-world environments. The theoretical foundations are examined in dialogue with, and in advance of, established co-training principles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。