提出Eureka框架与基准,打破模型单一排名,揭示各模型真实优劣。
Eureka: Evaluating and Understanding Large Foundation Models
- 构建开源评估框架,实现多维度能力对比而非单一得分排名
- 设计Eureka-Bench基准,聚焦语言与多模态中被忽视的基础能力
- 分析12个主流模型,发现无绝对最优,不同模型各有专长
严谨可复现的评估对人工智能发展至关重要。由于基准饱和、方法不透明、生成任务测量困难以及模型能力覆盖面广等原因,评估面临挑战。本文提出三项贡献:首先,推出Eureka——一个开源框架,用于标准化大基础模型的评估,超越单一分数报告和排名;其次,引入Eureka-Bench,一个可扩展的基准集合,测试当前顶尖模型仍难以应对且被忽视的语言与多模态核心能力;其非饱和特性使我们能在能力层面发现模型间有意义的差异。第三,利用Eureka分析12个前沿模型,深入揭示模型失败原因与对比结果,为针对性改进提供依据。与近期榜单宣称某模型绝对领先不同,本研究显示并无唯一最优模型,不同模型在特定能力上表现更优。尽管已有进步,当前模型在图像细节理解、多模态输入依赖性、事实性与信息检索的准确性及拒绝响应过度等问题上仍存在显著不足。
原文摘要 · Abstract (English)
Rigorous and reproducible evaluation is critical for assessing the state of the art and for guiding scientific advances in Artificial Intelligence. Evaluation is challenging in practice due to several reasons, including benchmark saturation, lack of transparency in methods used for measurement, development challenges in extracting measurements for generative tasks, and, more generally, the extensive number of capabilities required for a well-rounded comparison across models. We make three contributions to alleviate the above challenges. First, we present Eureka, an open-source framework for standardizing evaluations of large foundation models beyond single-score reporting and rankings. Second, we introduce Eureka-Bench as an extensible collection of benchmarks testing capabilities that (i) are still challenging for state-of-the-art models and (ii) represent fundamental but overlooked language and multimodal capabilities. The inherent space for improvement in non-saturated benchmarks enables us to discover meaningful differences between models at a capability level. Third, using Eureka, we conduct an analysis of 12 state-of-the-art models, providing in-depth insights into failure understanding and model comparison, which can be leveraged to plan targeted improvements. In contrast to recent trends in reports and leaderboards showing absolute rankings and claims for one model or another to be the best, our analysis shows that there is no such best model. Different models have different strengths, but there are models that appear more often than others as best performers for some capabilities. Despite the recent improvements, current models still struggle with several fundamental capabilities including detailed image understanding, benefiting from multimodal input when available rather than fully relying on language, factuality and grounding for information retrieval, and over refusals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。