提出首个支持多模态相关性评估的新指标MAJORScore
MAJORScore: A Novel Metric for Evaluating Multimodal Relevance via Joint Representation
- 基于多模态联合表示,统一不同模态的潜在空间
- 对一致模态评分提升26.03%~64.29%,不一致时下降13.28%~20.54%
- 适合大规模多模态数据与模型性能评估
现有跨模态相关性评估通常依赖预训练对比学习模型(如CLIP)的嵌入能力,仅适用于双模态分析,难以评估多模态相似性。本文首次提出MAJORScore,一种基于多模态联合表示的新评估指标,可处理N个模态(N≥3)的相关性计算。通过将多模态数据映射到同一潜在空间,实现跨模态的统一尺度表征,支持公平的相关性评分。大量实验表明,相比现有方法,MAJORScore在模态一致性场景下提升26.03%~64.29%,在不一致场景下降低13.28%~20.54%。该指标为大规模多模态数据集及多模态模型性能评估提供了更可靠的依据。
原文摘要 · Abstract (English)
The multimodal relevance metric is usually borrowed from the embedding ability of pretrained contrastive learning models for bimodal data, which is used to evaluate the correlation between cross-modal data (e.g., CLIP). However, the commonly used evaluation metrics are only suitable for the associated analysis between two modalities, which greatly limits the evaluation of multimodal similarity. Herein, we propose MAJORScore, a brand-new evaluation metric for the relevance of multiple modalities ($N$ modalities, $N\ge3$) via multimodal joint representation for the first time. The ability of multimodal joint representation to integrate multiple modalities into the same latent space can accurately represent different modalities at one scale, providing support for fair relevance scoring. Extensive experiments have shown that MAJORScore increases by 26.03%-64.29% for consistent modality and decreases by 13.28%-20.54% for inconsistence compared to existing methods. MAJORScore serves as a more reliable metric for evaluating similarity on large-scale multimodal datasets and multimodal model performance evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。