arXiv:2606.22723cs.CL2026-06

评测大模型在巴西大学入学考开放题的表现,覆盖数学与图像理解难点。

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams

论文配图:BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams
图 1 · 摘自论文原文
  • 基于两所顶尖巴西大学的开放问答题构建新基准
  • 21个模型表现差距达4.92分,数学与图像理解最弱
  • 含图文数据,适合评估多模态和生成能力

尽管大语言模型在诸多任务中表现优异,但针对葡萄牙语的评估仍不足,尤其缺乏对需要深度推理与自由生成的开放性论述题的考察。原有BLUEX基准虽通过巴西大学入学考试的单选题填补了葡萄牙语数据空白,却未涵盖更难的第二阶段考试——该阶段要求自由作答。本文提出BLUEX v2,源自巴西两所顶尖高校UNICAMP(Comvest)和USP(Fuvest)2022–2025年的第二阶段试题,包含395道题,衍生出919个评分子题,其中55.7%的问题附带图像(推理时以情境感知描述呈现,支持视觉与纯文本模型统一评估)。每道题标注学科领域、官方参考答案、由大模型生成的评分标准及六类认知能力标签。我们采用大模型作为评判者协议,评估21个先进模型,结果表明模型性能跨度达4.92分(0–10分制,4.18–9.10),数学推理与图像理解为最薄弱环节。评估代码、模型输出与数据集已公开于GitHub与Hugging Face。

原文摘要 · Abstract (English)

Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended, discursive tasks that demand deeper reasoning and generation capabilities. While the original BLUEX benchmark addressed the scarcity of Portuguese evaluation datasets through multiple-choice questions from Brazilian university entrance exams, it did not cover the more challenging second-phase examinations, which require free-form written responses. In this work, we introduce BLUEX v2, a benchmark derived from the second-phase entrance exams of Brazil's two leading universities: UNICAMP (Comvest) and USP (Fuvest), spanning exam years 2022--2025. Our dataset comprises 395 questions unfolding into 919 graded subquestions, with 55.7% of questions containing associated images (represented as context-aware captions during inference to enable evaluation across both vision-capable and text-only models). Each question is annotated with subject area, official reference answers, LLM-generated rubric criteria, and six cognitive capability tags. We evaluate 21 state-of-the-art LLMs using an LLM-as-a-judge protocol. Results reveal a 4.92-point performance spread across models (4.18-9.10 on a 0-10 scale), with Mathematical Reasoning and Image Understanding emerging as the hardest capability dimensions. The evaluation code, model outputs, and dataset are publicly available at https://github.com/TropicAI-Research/BLUEXv2 and on Hugging Face at https://huggingface.co/datasets/Tropic-AI/BLUEX-v2.

大模型评测葡萄牙语开放问答多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。