解决多模态数据长尾分布问题,提升少数类识别准确率。
Simultaneous Long-tailed Recognition and Multi-modal Fusion for Highly Imbalanced Multi-modal Data

- 用动态加权融合多模态信息,优先利用更可靠的模态。
- 在真实数据集上,对少数类的识别精度提升显著。
- 适合处理图像与表格数据混合的不平衡场景。
类别不平衡数据中的长尾分布是深度学习模型面临的根本挑战,模型易偏向多数类。尽管近期长尾识别方法已有所缓解,但大多局限于单模态输入,无法充分利用多种数据源的互补信息。本文提出一种新型长尾识别框架,显式处理多模态输入。通过将异构数据融合为统一表示,并利用模态特异性网络估计各模态的可信度,以信心引导的权重动态调节融合过程,使信息量更高的模态对最终决策贡献更大。为进一步提升性能,设计了适配多样化模态组合(如图像与表格数据)的专用训练与测试流程。在基准与真实世界数据集上的大量实验表明,该方法不仅能有效整合多模态信息,还在长尾、类别不平衡场景下优于现有方法,展现出强鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Long-tailed distributions in class-imbalanced data present a fundamental challenge for deep learning models, which tend to be biased toward majority classes. While recent methods for long-tailed recognition have mitigated this issue, they are largely restricted to single-modal inputs and cannot fully exploit complementary information from diverse data sources. In this work, we introduce a new framework for long-tailed recognition that explicitly handles multi-modal inputs. Our approach extends multi-expert architectures to the multi-modal setting by fusing heterogeneous data into a unified representation while leveraging modality-specific networks to estimate the informativeness of each modality. These confidence-guided weights dynamically modulate the fusion process, ensuring that more informative modalities contribute more strongly to the final decision. To further enhance performance, we design specialized training and test procedures that accommodate diverse modality combinations, including images and tabular data. Extensive experiments on benchmark and real-world datasets demonstrate that the proposed approach not only effectively integrates multi-modal information but also outperforms existing methods in handling long-tailed, class-imbalanced scenarios, highlighting its robustness and generalization capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。