评测大模型解极值问题能力,发现现有数学测试不全面
Max It or Miss It: Benchmarking LLM On Solving Extremal Problems
- 构建极值求解基准数据集ExtremBench,含93个标准化题目
- 多款主流模型在极值求解上表现参差,与传统数学测试结果不一致
- 适合研究模型推理能力评估或数学教育应用的开发者参考
测试时扩展已使大语言模型在数学推理领域展现出卓越能力,尤其通过中间链式思维(CoT)推理生成最终答案。然而,这些推理能力的具体来源与机制仍不清晰。优化推理,即在约束条件下寻找极值,是规划、控制、资源分配和提示搜索等关键应用的基本抽象。为系统评估该能力,我们引入ExtremBench,一个从中国数学奥林匹克不等式题改编而来的极值求解基准数据集,包含93个标准化极值问题。我们在多个领先开源模型家族(包括Qwen3、GPT-OSS和DeepSeek)上进行了广泛评估。结果表明,大模型的极值求解推理能力并不总与现有数学基准(如AIME25和MATH-500)一致,部分模型虽具强大泛化数学推理能力但极值求解表现差,反之亦然。这一差异凸显当前评估方法的关键缺口,暗示现有基准可能无法全面覆盖数学推理的全部能力谱系。
原文摘要 · Abstract (English)
Test-time scaling has enabled Large Language Models (LLMs) with remarkable reasoning capabilities, particularly in mathematical domains, through intermediate chain-of-thought (CoT) reasoning before generating final answers. However, the specific sources and mechanisms underlying these reasoning capabilities remain insufficiently understood. Optimization reasoning, i.e. finding extrema under constraints, represents a fundamental abstraction that underpins critical applications in planning, control, resource allocation, and prompt search. To systematically evaluate this capability, we introduce ExtremBench, a benchmark dataset for solving mathematical extremal problems, curated from inequality exercises used for Chinese Mathematical Olympiad and transformed into $93$ standardized extrema-finding problems. We conduct extensive evaluations across various state-of-the-art open-source model families, including the Qwen3, GPT-OSS, and DeepSeek. Our results reveal that LLMs' extremal-solving reasoning capabilities do not always align with those of current mathematical benchmarks such as AIME25 and MATH-500, with some models showing strong general mathematical reasoning but poor extremal-solving skills, and vice versa. This discrepancy highlights a critical gap in current evaluation practices and suggests that existing benchmarks may not comprehensively capture the full spectrum of mathematical reasoning abilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。