arXiv:2504.04310cs.CLcs.AI2025-04AAAI被引 37

构建36个真实组合优化问题的基准,评测大模型智能体的算法搜索能力。

CO-Bench: Benchmarking Language Model Agents in Algorithm Search for Combinatorial Optimization

  • 设计包含36个真实问题的结构化基准,覆盖多领域复杂度。
  • 对比大模型智能体与人工设计算法,发现其在复杂约束下的局限性。
  • 适合研究大模型用于算法自动设计的研究者和开发者。

尽管基于大语言模型(LLM)的智能体在软件工程和机器学习研究等领域受到广泛关注,其在组合优化(CO)中的作用仍相对未被充分探索。这一差距凸显了深入理解其解决结构化、约束密集型问题潜力的必要性,而当前研究受限于缺乏全面的基准体系。为此,我们提出CO-Bench,一个涵盖36个来自多个领域和复杂度级别的真实世界组合优化问题的基准套件。该基准提供结构化问题定义和精选数据,支持对LLM智能体的严谨评估。通过将多种智能体框架与已有的人工设计算法进行对比,我们揭示了现有LLM智能体的优势与局限,并指明未来研究的可行方向。CO-Bench已公开发布于https://github.com/sunnweiwei/CO-Bench。

原文摘要 · Abstract (English)

Although LLM-based agents have attracted significant attention in domains such as software engineering and machine learning research, their role in advancing combinatorial optimization (CO) remains relatively underexplored. This gap underscores the need for a deeper understanding of their potential in tackling structured, constraint-intensive problems -- a pursuit currently limited by the absence of comprehensive benchmarks for systematic investigation. To address this, we introduce CO-Bench, a benchmark suite featuring 36 real-world CO problems drawn from a broad range of domains and complexity levels. CO-Bench includes structured problem formulations and curated data to support rigorous investigation of LLM agents. We evaluate multiple agentic frameworks against established human-designed algorithms, revealing the strengths and limitations of existing LLM agents and identifying promising directions for future research. CO-Bench is publicly available at https://github.com/sunnweiwei/CO-Bench.

组合优化大模型智能体基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。