arXiv:2511.18649cs.CL2025-11被引 1

用2026年韩国高考数学题测试大模型,确保数据零泄露。

Evaluating Large Language Models on the 2026 Korean CSAT Mathematics Exam: Measuring Mathematical Ability in a Zero-Data-Leakage Setting

  • 在无数据泄露环境下,用真实高考题评估24个大模型数学能力。
  • GPT-5等模型达100分,小模型gpt-oss-20B也超95分。
  • 文本输入优于图像,高难度微积分题仍是短板,推理越强越耗资源。

本研究利用2026年韩国大学入学能力考试(CSAT)数学部分,系统评估大型语言模型(LLMs)的数学推理能力,确保完全无数据泄露的评测环境。为解决现有基准的数据泄露问题,我们在考试公开后两小时内数字化了全部46道题(22道共同题、24道选考题),彻底排除其进入模型训练数据的可能性。对24个前沿LLMs进行了全面评估,涵盖文本、图像、图文混合输入模态及韩语、英语提示语言。GPT-5系列在特定配置下取得满分100分,Grok 4、Qwen 3 235B和Gemini 2.5 Pro得分均超过97分。值得注意的是,尽管规模较小,gpt-oss-20B仍取得95.7分,表现出高性价比。问题层面分析显示,微积分是表现最弱的领域,尤其在4分高难度题上性能显著下降。文本输入始终优于图像输入,而提示语言的影响因模型规模而异。在对GPT-5系列的推理增强实验中,增加推理强度使得分从82.6提升至100,但令牌消耗增至四倍,效率大幅降低,表明轻量级推理模型可能更具实用性。本研究贡献包括:(1)实现完全无暴露的评测环境;(2)建立标准化的数字化流程,将人类导向的考试材料转化为适用于LLM的评测数据;(3)提出融合性能、成本与时间考量的实用评估视角。详细结果与模型对比可在2026韩国CSAT LLM评估排行榜查看:https://isoft.cnu.ac.kr/csat2026/

原文摘要 · Abstract (English)

This study systematically evaluated the mathematical reasoning capabilities of Large Language Models (LLMs) using the 2026 Korean College Scholastic Ability Test (CSAT) Mathematics section, ensuring a completely contamination-free evaluation environment. To address data leakage issues in existing benchmarks, we digitized all 46 questions (22 common and 24 elective) within two hours of the exam's public release, eliminating any possibility of inclusion in model training data. We conducted comprehensive evaluations of 24 state-of-the-art LLMs across varying input modalities (Text-only, Image-only, Text+Figure) and prompt languages (Korean, English). The GPT-5 family models achieved perfect scores (100 points) under a limited set of language-modality configurations, while Grok 4, Qwen 3 235B, and Gemini 2.5 pro also scored above 97 points. Notably, gpt-oss-20B achieved 95.7 points despite its relatively small size, demonstrating high cost-effectiveness. Problem-specific analysis revealed Calculus as the weakest domain with significant performance degradation on 4-point high-difficulty problems. Text input consistently outperformed image input, while prompt language effects varied by model scale. In reasoning enhancement experiments with GPT-5 series, increased reasoning intensity improved performance (82.6->100 points) but quadrupled token usage and drastically reduced efficiency, suggesting that models with minimal reasoning may be more practical. This research contributes: (1) implementation of a completely unexposed evaluation environment, (2) a standardized digitization pipeline that converts human-targeted exam materials into LLM-ready evaluation data, and (3) a practical evaluation perspective integrating performance, cost, and time considerations. Detailed results and model comparisons are available at the 2026 Korean CSAT LLM Evaluation Leaderboard; https://isoft.cnu.ac.kr/csat2026/

大模型评测数学推理高考题零数据泄露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。