GAUSS benchmark全方位评估大模型数学能力,揭示其真实水平。
GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models
- 按认知能力分12个维度,细粒度评估数学技能
- 构建可解释的模型数学能力画像,反映真实智能水平
- 适合研究者对比模型优劣,指导模型优化
我们提出 extbf{GAUSS}( extbf{G}eneral extbf{A}ssessment of extbf{U}nderlying extbf{S}tructured extbf{S}kills in Mathematics),一个涵盖十二项核心数学技能维度的基准测试,分为知识与理解、问题求解与表达、元技能与创造力三个领域。通过按认知能力分类题目并设计隔离特定能力的任务,GAUSS构建了全面、精细且可解释的模型数学能力画像,准确反映其底层数学智能。为展示使用方法,我们分析了 extsc{GPT-5-thinking}的技能表现,揭示其优势与短板,并与 extsc{o4-mini-high}对比,凸显多维技能评估的价值。
原文摘要 · Abstract (English)
We introduce \textbf{GAUSS} (\textbf{G}eneral \textbf{A}ssessment of \textbf{U}nderlying \textbf{S}tructured \textbf{S}kills in Mathematics), a benchmark that evaluates LLMs' mathematical abilities across twelve core skill dimensions, grouped into three domains: knowledge and understanding, problem solving and communication, and meta-skills and creativity. By categorizing problems according to cognitive skills and designing tasks that isolate specific abilities, GAUSS constructs comprehensive, fine-grained, and interpretable profiles of models' mathematical abilities. These profiles faithfully represent their underlying mathematical intelligence. To exemplify how to use the \textsc{GAUSS} benchmark, we have derived the skill profile of \textsc{GPT-5-thinking}, revealing its strengths and weaknesses as well as its differences relative to \textsc{o4-mini-high}, thereby underscoring the value of multidimensional, skill-based evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。