用顶级编程竞赛题评估大模型,更真实反映代码能力差距。
AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions
- 基于IOI、ICPC等顶级竞赛题设计高难度评测集。
- 引入专家验证的完整测试用例,避免低质量数据干扰。
- 适合研究代码推理与真实编程能力评估的学者参考。
编程竞赛已成为评估大语言模型(LLMs)推理与编码能力的关键基准。尽管现有评测取得显著进展,但当前评估夸大了模型的实际水平,掩盖了其与顶尖人类程序员之间的巨大差距。这一差距源于两个核心问题:基准题目的难度和覆盖范围不足,以及由低质量测试用例引发的评估偏差。为此,我们提出AetherCode,一个源自IOI、ICPC等顶级编程竞赛的新基准,具备更广覆盖与更高难度。AetherCode结合自动化生成与人工校验,构建了全面且经过专家验证的测试用例集,确保评估的严谨性与可靠性。通过难题设计与严格评测相结合,AetherCode为衡量LLM的真实能力提供了更准确的尺度,并为未来代码推理研究设立了新标准。
原文摘要 · Abstract (English)
Competitive programming has emerged as a critical benchmark for evaluating the reasoning and coding capabilities of Large Language Models (LLMs). Despite impressive progress on existing benchmarks, we argue that current evaluations overstate model proficiency, masking a substantial gap between LLMs and elite human programmers. This gap arises from two key limitations: insufficient difficulty and scope of benchmark problems, and evaluation bias from low-quality test cases. To address these shortcomings, we present AetherCode, a new benchmark that draws problems from premier programming competitions such as IOI and ICPC, offering broader coverage and higher difficulty. AetherCode further incorporates comprehensive, expert-validated test suites built through a hybrid of automated generation and human curation, ensuring rigorous and reliable assessment. By combining challenging problem design with robust evaluation, AetherCode provides a more faithful measure of LLM capabilities and sets a new standard for future research in code reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。