构建跨领域多模态融合评估基准,解决模型评价碎片化问题
MULTIBENCH++: A Unified and Comprehensive Multimodal Fusion Benchmarking Across Specialized Domains
- 整合30+数据集、15种模态和20项任务,覆盖多领域真实场景
- 建立统一自动化评估流水线,实现主流模型与融合方法的公平对比
- 提供可复现基准,助力通用多模态模型研发
尽管多模态融合取得显著进展,其发展仍受限于缺乏充分的评估基准。当前方法通常仅在少数公开数据集上测试,范围有限,无法充分反映现实场景的复杂性与多样性,可能导致评价偏差。这带来双重挑战:一方面,模型可能过拟合特定数据集的偏见,影响在更广泛实际应用中的泛化能力;另一方面,缺乏统一评估标准,使得不同融合方法间的公平比较困难。为此,我们构建了一个大规模、领域自适应的多模态评估基准,整合超过30个数据集,涵盖15种模态和20项预测任务,覆盖关键应用领域。同时,开发了开源、统一、自动化的评估流水线,包含先进模型的标准实现与多样融合范式。基于该平台,我们开展了大规模实验,成功建立了多个任务的新性能基准。本工作为学术界提供了严谨可复现的多模态模型评估平台,旨在推动多模态人工智能迈向新高度。
原文摘要 · Abstract (English)
Although multimodal fusion has made significant progress, its advancement is severely hindered by the lack of adequate evaluation benchmarks. Current fusion methods are typically evaluated on a small selection of public datasets, a limited scope that inadequately represents the complexity and diversity of real-world scenarios, potentially leading to biased evaluations. This issue presents a twofold challenge. On one hand, models may overfit to the biases of specific datasets, hindering their generalization to broader practical applications. On the other hand, the absence of a unified evaluation standard makes fair and objective comparisons between different fusion methods difficult. Consequently, a truly universal and high-performance fusion model has yet to emerge. To address these challenges, we have developed a large-scale, domain-adaptive benchmark for multimodal evaluation. This benchmark integrates over 30 datasets, encompassing 15 modalities and 20 predictive tasks across key application domains. To complement this, we have also developed an open-source, unified, and automated evaluation pipeline that includes standardized implementations of state-of-the-art models and diverse fusion paradigms. Leveraging this platform, we have conducted large-scale experiments, successfully establishing new performance baselines across multiple tasks. This work provides the academic community with a crucial platform for rigorous and reproducible assessment of multimodal models, aiming to propel the field of multimodal artificial intelligence to new heights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。