arXiv:2508.15835cs.CLcs.AI2025-08被引 1

评测大模型在巴西高考中的表现,发现数学和工程题仍有短板。

Alvorada-Bench: Can Language Models Solve Brazilian University Entrance Exams?

  • 构建4515题巴西高考文本基准,覆盖五所名校真题。
  • 顶尖模型总准确率超94%,但数学与工程类题目表现下降。
  • 模型自信心评估准确,且低成本即可实现高精度。

语言模型在巴西的应用日益广泛,但现有评估多以英语为主。本文提出 Alvorada-Bench,一个包含4,515道题的纯文本基准,源自五所巴西大学入学考试。对20个模型进行零样本、角色扮演和思维链提示测试,生成270,900条带结构化自评(置信度、感知难度、Bloom层级)的回答。顶级模型整体准确率超过94%,但在数学及侧重工程的IME和ITA考试中准确率下降,表明多步推理仍存缺陷。模型置信度与感知难度高度相关,且校准良好,说明其可自我评估能力。成本-精度分析显示,每千次调用低于2美元即可实现高准确率。在2024年ENEM考试中,最优模型(O3)在语言科目取得满分,而最弱系统(GPT-4.1 Nano)仅在数学上逊于人类。通过整合数十年巴西教育重点与每年百万考生的选拔标准,Alvorada-Bench验证了大模型能否跨越语言、文化与推理交织的学术准备门槛。

原文摘要 · Abstract (English)

Language models are increasingly used in Brazil, but most evaluation remains English-centric. This paper presents Alvorada-Bench, a 4,515-question, text-only benchmark drawn from five Brazilian university entrance examinations. Evaluating twenty models under zero-shot, role-playing, and chain-of-thought prompting, producing 270,900 responses with structured self-reports of confidence, perceived difficulty, and Bloom level. The top models exceed 94% accuracy overall, but accuracy declines on Mathematics and on the engineering oriented IME and ITA exams, indicating persistent weaknesses in multi-step reasoning. Confidence is well calibrated and correlates with perceived difficulty, revealing that models can accurately assess their own certainty capabilities. A cost accuracy analysis shows that high accuracy is achievable at under $2 per 1K tokens. On ENEM 2024 the top model (O3) achieved perfect scores in Languages subject questions while even the weakest system (GPT-4.1 Nano) only underperforms humans in Mathematics. Through exams that distill decades of Brazilian educational priorities and assess millions of students yearly, Alvorada-Bench establishes whether language models can navigate the intersection of language, culture, and reasoning that defines academic readiness in Brazil.

大模型评测巴西高考多步推理语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。