arXiv:2504.20183cs.SEcs.AI2025-04中稿 · GECCO Workshop 202…被引 16

构建评估大模型自动设计优化算法的标准化测试框架。

BLADE: Benchmark suite for LLM-driven Automated Design and Evolution of iterative optimisation heuristics

  • 提出可扩展的基准测试框架BLADE,支持黑箱优化场景
  • 集成多个测试问题集与生成器,支持泛化、特化等能力评估
  • 适合研究大模型自动生成算法的学者和开发者

大语言模型(LLMs)在自动化算法发现(AAD)中的应用,特别是在优化启发式算法方面,正成为新兴研究方向。然而,现有基准测试存在局限性,且算法生成过程不透明,亟需更稳健、标准化的评估方法。为此,我们提出BLADE(LLM驱动的自动化设计与进化基准套件),一个专为连续黑箱优化场景设计的模块化、可扩展框架。BLADE整合了多个基准问题集(如MA-BBOB和SBOX-COST),配备实例生成器与文本描述,支持对泛化、特化及信息利用等能力的针对性测试。框架提供灵活实验配置、标准化日志记录以保障可复现性与公平比较,并内置代码演化图谱、可视化工具等分析方法,同时通过集成IOHanalyser和IOHexplainer,实现与人工设计基线的对比。通过两个应用场景验证了其有效性,分别探索突变提示策略与目标函数特化。该框架为系统评估LLM驱动的AAD方法提供了开箱即用的解决方案。

原文摘要 · Abstract (English)

The application of Large Language Models (LLMs) for Automated Algorithm Discovery (AAD), particularly for optimisation heuristics, is an emerging field of research. This emergence necessitates robust, standardised benchmarking practices to rigorously evaluate the capabilities and limitations of LLM-driven AAD methods and the resulting generated algorithms, especially given the opacity of their design process and known issues with existing benchmarks. To address this need, we introduce BLADE (Benchmark suite for LLM-driven Automated Design and Evolution), a modular and extensible framework specifically designed for benchmarking LLM-driven AAD methods in a continuous black-box optimisation context. BLADE integrates collections of benchmark problems (including MA-BBOB and SBOX-COST among others) with instance generators and textual descriptions aimed at capability-focused testing, such as generalisation, specialisation and information exploitation. It offers flexible experimental setup options, standardised logging for reproducibility and fair comparison, incorporates methods for analysing the AAD process (e.g., Code Evolution Graphs and various visualisation approaches) and facilitates comparison against human-designed baselines through integration with established tools like IOHanalyser and IOHexplainer. BLADE provides an `out-of-the-box' solution to systematically evaluate LLM-driven AAD approaches. The framework is demonstrated through two distinct use cases exploring mutation prompt strategies and function specialisation.

自动化设计优化算法大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。