评测AI解决复杂优化问题的长期迭代能力,聚焦真实工业场景。
ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering
- 基于真实竞赛题构建长周期优化任务,支持交互式反馈与可视化。
- 前沿大模型在单个问题上表现优异,但跨问题一致性与长期求解能力不足。
- 适合研究长时序决策、算法工程与强化学习的开发者与研究人员。
AI系统在解决包装配送路线、人员排班、工厂生产计划和电网平衡等复杂优化问题时表现如何?我们提出ALE-Bench,一个用于评估AI在基于得分的算法编程竞赛中表现的新基准。该基准源自AtCoder启发式竞赛的真实任务,包含计算上困难且无已知精确解的优化问题。与短时、通过/失败的编码评测不同,ALE-Bench鼓励在长时间跨度内对解决方案进行迭代优化。我们的软件框架支持交互式智能体架构,可利用测试运行反馈与可视化信息。对前沿大语言模型的评估显示,尽管它们在特定问题上表现良好,但在跨问题一致性及长期求解能力方面仍显著落后于人类。这凸显了该基准在推动未来AI发展中的必要性。
原文摘要 · Abstract (English)
How well do AI systems perform in algorithm engineering for hard optimization problems in domains such as package-delivery routing, crew scheduling, factory production planning, and power-grid balancing? We introduce ALE-Bench, a new benchmark for evaluating AI systems on score-based algorithmic programming contests. Drawing on real tasks from the AtCoder Heuristic Contests, ALE-Bench presents optimization problems that are computationally hard and admit no known exact solution. Unlike short-duration, pass/fail coding benchmarks, ALE-Bench encourages iterative solution refinement over long time horizons. Our software framework supports interactive agent architectures that leverage test-run feedback and visualizations. Our evaluation of frontier LLMs revealed that while they demonstrate high performance on specific problems, a notable gap remains compared to humans in terms of consistency across problems and long-horizon problem-solving capabilities. This highlights the need for this benchmark to foster future AI advancements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。