arXiv:2410.13754cs.AIcs.LG2024-10被引 6

构建首个跨模态真实数据评估基准,统一多模态模型评测标准。

MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures

  • 设计多模态混合数据与自适应修正流程,还原真实任务分布。
  • 模型排名与人工评估高度一致(相关性达0.98),效率更高。
  • 适合研究多模态评测与模型性能验证的学者使用。

感知与生成多样化模态对人工智能模型有效学习和交互现实信号至关重要,因此需要可靠的评估方法。当前评估存在两大问题:(1) 标准不一致,不同研究社区采用不同协议且成熟度各异;(2) 存在显著的查询、评分与泛化偏差。为此,我们提出 MixEval-X,首个面向任意输入输出模态的真实世界基准,旨在优化并标准化跨模态评估。通过构建多模态基准混合体与适应-修正流水线,重建真实任务分布,确保评估结果可有效泛化至真实应用场景。大规模元评估表明,该方法能有效对齐基准样本与真实任务分布。同时,MixEval-X 的模型排名与众包式真实世界评估高度一致(最高相关性达0.98),且效率显著提升。我们提供全面排行榜,重新排序现有模型与机构,并为理解多模态评估提供洞见,推动未来研究发展。

原文摘要 · Abstract (English)

Perceiving and generating diverse modalities are crucial for AI models to effectively learn from and engage with real-world signals, necessitating reliable evaluations for their development. We identify two major issues in current evaluations: (1) inconsistent standards, shaped by different communities with varying protocols and maturity levels; and (2) significant query, grading, and generalization biases. To address these, we introduce MixEval-X, the first any-to-any, real-world benchmark designed to optimize and standardize evaluations across diverse input and output modalities. We propose multi-modal benchmark mixture and adaptation-rectification pipelines to reconstruct real-world task distributions, ensuring evaluations generalize effectively to real-world use cases. Extensive meta-evaluations show our approach effectively aligns benchmark samples with real-world task distributions. Meanwhile, MixEval-X's model rankings correlate strongly with that of crowd-sourced real-world evaluations (up to 0.98) while being much more efficient. We provide comprehensive leaderboards to rerank existing models and organizations and offer insights to enhance understanding of multi-modal evaluations and inform future research.

多模态评估真实数据基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。