复杂多模态模型未必更优,简单基线反而更可靠。
Fusion or Confusion? Multimodal Complexity Is Not All You Need
- 复现19个顶尖方法,在9个数据集上对比复杂模型与简单基线。
- 多数复杂模型表现不如简单基线,且结果不一致。
- 适合关注方法严谨性而非架构创新的研究者。
多模态学习已成为热门领域,理论上融合多源信息可显著提升性能。然而,当前模型趋向于复杂的深度架构,普遍假设专用方法能带来性能提升。我们通过大规模实证研究,复现了19个高影响力多模态方法,覆盖九个多样化数据集,最多涉及23种模态。在标准化实验条件下(包括超参数调优、权重初始化、交叉验证与统计检验),复杂模型常导致信息混淆,而非有效融合。结果表明,复杂架构并未稳定优于单模态基线或简单基线(SimBaMM)。进一步案例分析显示,顶级论文中也存在方法论缺陷,凸显标准化评估的必要性。我们主张:多模态研究应从追求架构新颖转向重视方法严谨性。
原文摘要 · Abstract (English)
Multimodal learning has become a prominent research area, with the potential of substantial performance gains by combining information across modalities. At the same time, model development has trended toward increasingly complex deep learning architectures, motivated by the assumption that multimodal-specific methods improve performance. We challenge this assumption through a large-scale empirical study by reimplementing 19 high-impact multimodal methods across nine diverse datasets with up to 23 modalities. Under standardized experimental conditions, including hyperparameter tuning, weight initialization, cross-validation, and statistical testing, increased multimodal complexity often yields confusion rather than effective fusion of data modalities. Accordingly, complex multimodal architectures do not reliably outperform unimodal baselines and a Simple Baseline for Multimodal Learning (SimBaMM). Through a focused case study, we further demonstrate concrete methodological shortcomings even in top-tier multimodal learning publications, underscoring the need for standardized evaluation practices. In summary, we argue for a shift in focus for multimodal learning: away from the pursuit of architectural novelty and toward methodological rigor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。