提出多模态泛化新任务,解决模型遇新模态时失效的问题
Towards Modality Generalization: A Benchmark and Prospective Analysis
- 区分弱泛化与强泛化两种场景,构建可评估的基准测试
- 现有方法在未见模态上性能显著下降,暴露泛化能力不足
- 适合研究多模态鲁棒性、跨模态迁移的学者参考
多模态学习通过融合多种模态信息,在识别与检索任务中表现优于单模态方法。然而,由于资源和隐私限制,真实场景中常出现训练时未见过的新模态,当前方法难以应对。本文提出多模态泛化(Modality Generalization, MG)任务,旨在让模型适应未见过的模态。定义两种情形:弱MG下,已知模态可通过现有感知器映射至联合嵌入空间;强MG下,不存在此类映射关系。为推动研究,我们构建了包含多模态算法的综合性基准,并适配现有泛化方法。大量实验揭示了MG的复杂性,暴露出现有方法的局限性,并指明未来关键研究方向。本工作为构建更鲁棒、可适应的多模态模型奠定了基础,使其能在真实场景中处理未见模态。
原文摘要 · Abstract (English)
Multi-modal learning has achieved remarkable success by integrating information from various modalities, achieving superior performance in tasks like recognition and retrieval compared to uni-modal approaches. However, real-world scenarios often present novel modalities that are unseen during training due to resource and privacy constraints, a challenge current methods struggle to address. This paper introduces Modality Generalization (MG), which focuses on enabling models to generalize to unseen modalities. We define two cases: Weak MG, where both seen and unseen modalities can be mapped into a joint embedding space via existing perceptors, and Strong MG, where no such mappings exist. To facilitate progress, we propose a comprehensive benchmark featuring multi-modal algorithms and adapt existing methods that focus on generalization. Extensive experiments highlight the complexity of MG, exposing the limitations of existing methods and identifying key directions for future research. Our work provides a foundation for advancing robust and adaptable multi-modal models, enabling them to handle unseen modalities in realistic scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。