构建多模态不平衡学习基准,系统评估主流方法优劣。
BalanceBenchmark: A Survey for Multimodal Imbalance Learning
- 按策略将主流方法分为四类,统一评估框架
- 涵盖多数据集与三维度指标,揭示性能与复杂度权衡
- 工具开源,适合研究者对比算法或设计新模型
多模态学习因能融合多源信息而备受关注,但常受多模态不平衡问题困扰——某些模态主导,其他模态被忽略。尽管已有多种缓解方法提出,但缺乏全面且公平的比较。本文系统地将主流多模态不平衡算法按缓解策略分为四类,并引入 BalanceBenchmark 基准,包含多个常用多维数据集及从性能、不平衡程度、复杂度三个角度的评估指标。为确保公平比较,我们开发了模块化可扩展的工具包,标准化不同方法的实验流程。基于该基准的实验揭示了各类方法在性能、平衡度与计算复杂度方面的特性与优势。本分析有望启发未来更高效的方法,以及基础模型的设计。工具代码已公开于 https://github.com/GeWu-Lab/BalanceBenchmark。
原文摘要 · Abstract (English)
Multimodal learning has gained attention for its capacity to integrate information from different modalities. However, it is often hindered by the multimodal imbalance problem, where certain modality dominates while others remain underutilized. Although recent studies have proposed various methods to alleviate this problem, they lack comprehensive and fair comparisons. In this paper, we systematically categorize various mainstream multimodal imbalance algorithms into four groups based on the strategies they employ to mitigate imbalance. To facilitate a comprehensive evaluation of these methods, we introduce BalanceBenchmark, a benchmark including multiple widely used multidimensional datasets and evaluation metrics from three perspectives: performance, imbalance degree, and complexity. To ensure fair comparisons, we have developed a modular and extensible toolkit that standardizes the experimental workflow across different methods. Based on the experiments using BalanceBenchmark, we have identified several key insights into the characteristics and advantages of different method groups in terms of performance, balance degree and computational complexity. We expect such analysis could inspire more efficient approaches to address the imbalance problem in the future, as well as foundation models. The code of the toolkit is available at https://github.com/GeWu-Lab/BalanceBenchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。