arXiv:2506.09745cs.CV2025-06

解决多模态数据类别不一致问题,提升跨模态识别能力。

Class Similarity-Based Multimodal Classification under Heterogeneous Category Sets

  • 通过类别相似性对齐不同模态特征到共享语义空间。
  • 基于不确定性选择关键模态,融合时用辅助模态优化主模态预测。
  • 适用于真实场景中类别分布不一致的多模态分类任务。

现有多模态方法通常假设各模态共享相同类别集,但现实中多模态数据的类别分布常不一致,影响模型利用跨模态信息识别所有类别的能力。本文提出实际设置——多模态异构类别集学习(MMHCL),即模型在异构类别集上训练,测试时需识别所有模态的完整类别集。为此,我们提出基于类别相似性的跨模态融合模型(CSCF):首先将模态特定特征对齐至共享语义空间,实现已见与未见类间知识迁移;其次通过不确定性估计选择最具判别力的模态进行决策融合;最后根据类别相似性整合跨模态信息,由辅助模态优化主导模态的预测结果。实验表明,该方法在多个基准数据集上显著优于现有SOTA方法,有效解决了MMHCL任务。

原文摘要 · Abstract (English)

Existing multimodal methods typically assume that different modalities share the same category set. However, in real-world applications, the category distributions in multimodal data exhibit inconsistencies, which can hinder the model's ability to effectively utilize cross-modal information for recognizing all categories. In this work, we propose the practical setting termed Multi-Modal Heterogeneous Category-set Learning (MMHCL), where models are trained in heterogeneous category sets of multi-modal data and aim to recognize complete classes set of all modalities during test. To effectively address this task, we propose a Class Similarity-based Cross-modal Fusion model (CSCF). Specifically, CSCF aligns modality-specific features to a shared semantic space to enable knowledge transfer between seen and unseen classes. It then selects the most discriminative modality for decision fusion through uncertainty estimation. Finally, it integrates cross-modal information based on class similarity, where the auxiliary modality refines the prediction of the dominant one. Experimental results show that our method significantly outperforms existing state-of-the-art (SOTA) approaches on multiple benchmark datasets, effectively addressing the MMHCL task.

多模态类别不一致跨模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。