arXiv:2503.06885cs.CV2025-03ACL

评测多模态大模型在专业级开放任务中的表现,发现其在视觉理解与推理上仍有明显短板。

ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks

  • 构建跨10大领域56子领域的专家级多模态任务数据集
  • 24个主流模型在该测试中均未达人类专家水平
  • 适合关注多模态AI评估与进化的研究者参考

解决专家级多模态任务是迈向通用智能的关键里程碑。随着多模态大语言模型(MLLMs)能力持续提升,对这类先进多模态智能的评估变得必要但极具挑战。本文提出ProBench,一个包含4,000个由专业人士根据日常生产需求独立提交的开放性用户查询基准,覆盖科学、艺术、人文、编程、数学和创意写作等10个领域及56个子领域。通过MLLM-as-a-Judge方法,我们对24个最新模型进行了实验评估。结果表明,尽管顶尖开源模型已接近商用模型表现,但ProBench在视觉感知、文本理解、领域知识和高级推理方面仍构成显著挑战,为未来多模态AI研究提供了重要方向。

原文摘要 · Abstract (English)

Solving expert-level multimodal tasks is a key milestone towards general intelligence. As the capabilities of multimodal large language models (MLLMs) continue to improve, evaluation of such advanced multimodal intelligence becomes necessary yet challenging. In this work, we introduce ProBench, a benchmark of open-ended user queries that require professional expertise and advanced reasoning. ProBench consists of 4,000 high-quality samples independently submitted by professionals based on their daily productivity demands. It spans across 10 fields and 56 sub-fields, including science, arts, humanities, coding, mathematics, and creative writing. Experimentally, we evaluate and compare 24 latest models using MLLM-as-a-Judge. Our results reveal that although the best open-source models rival the proprietary ones, ProBench presents significant challenges in visual perception, textual understanding, domain knowledge and advanced reasoning, thus providing valuable directions for future multimodal AI research efforts.

多模态评测基准大模型专家任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。