用先进模型自动生成图像描述,让中文评测数据更全面。
BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
- 用最新大模型自动补全图像描述,提升数据可用性
- 新增1422个可使用问题,数量翻倍于原版
- 适合研究多语言评测与模型视觉理解的开发者
随着大语言模型能力的提升,亟需在多语言及非英语语境下建立稳健的评估方法。我们更新了BLUEX数据集,新增2024-2025年考试内容,并利用前沿模型自动生成图像描述,增强其在大模型预训练数据污染研究中的适用性。通过描述生成策略,使文本模型的可访问性提升超过40%,产生1,422个可用问题,数量超过原始BLUEX的两倍。我们评估了商用与开源大模型利用图像描述进行视觉上下文理解的能力。
原文摘要 · Abstract (English)
With the growing capabilities of Large Language Models (LLMs), there is an increasing need for robust evaluation methods, especially in multilingual and non-English contexts. We present an updated version of the BLUEX dataset, now including 2024-2025 exams and automatically generated image captions using state-of-the-art models, enhancing its relevance for data contamination studies in LLM pretraining. Captioning strategies increase accessibility to text-only models by more than 40%, producing 1,422 usable questions, more than doubling the number in the original BLUEX. We evaluated commercial and open-source LLMs and their ability to leverage visual context through captions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。