arXiv:2505.19614cs.LGcs.CV2025-05被引 6

多模态学习中的多重映射是固有挑战,忽略它会导致模型训练不稳、评估不可靠。

Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning

  • 指出多模态关系本质是多对多,非传统的一对一假设
  • 证明忽略多重性会引发训练不确定性与数据质量下降
  • 呼吁构建考虑多重性的新框架与评测标准,适合多模态研究者

多模态学习虽取得显著进展,尤其在跨模态大规模预训练方面,但当前方法大多基于模态间确定性一一对应假设。然而,真实世界中的多模态关系本质上是多对多的。这种多重性(multiplicity)并非噪声或标注错误所致,而是模态内变异性、表征不对称性和任务依赖模糊性的必然结果。本文认为,多重性是影响多模态学习全链条的根本瓶颈:从数据构建到模型训练再到评估基准。通过形式化其成因与后果,我们证明忽略多重性将导致训练不确定性、不可靠评估和数据质量下降。本文呼吁发展多重性感知的学习框架,以及新的数据集构建与评估协议。

原文摘要 · Abstract (English)

Multimodal learning has seen remarkable progress, particularly with large-scale pre-training across various modalities. Most current approaches are built on the assumption of a deterministic one-to-one alignment between modalities. However, this oversimplifies real-world multimodal relationships, where their nature is inherently many-to-many. The many-to-many property, or multiplicity, is not a side-effect of noise or annotation error, but an inevitable outcome of intra-modal variability, representational asymmetry, and task-dependent ambiguity in multimodal tasks. We argue that multiplicity is a fundamental bottleneck that affects all stages of the multimodal learning pipeline: from data construction to model training and evaluation benchmarks. By formalizing its causes and consequences, we demonstrate how ignoring multiplicity leads to training uncertainty, unreliable evaluation, and degraded dataset quality. This position paper calls for new research directions on multimodal learning, including multiplicity-aware learning frameworks and dataset construction and evaluation protocols.

多模态学习多重性数据构建评估标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。