arXiv:2601.12723cs.NEcs.AI2026-01

用大模型自动生成能区分算法优劣的优化测试题

An Evolutionary Framework for Automatic Optimization Benchmark Generation via Large Language Models

  • 让大模型当进化算子,自动生成数学表达式形式的优化题
  • 生成的问题中目标算法在80%以上实验里表现更优
  • 可生成对算法特性敏感的难题,适合算法评测与分析

优化基准在评估算法性能中起基础作用;然而现有人工基准难以捕捉真实问题结构的多样性和不规则性,而基于真实问题的基准又成本高、难构建。为此,我们提出一种基于大语言模型(LLM)的进化式自动基准生成框架,称为LLM驱动的进化基准生成器(LLM-EBG)。该框架将LLM作为进化算子,在灵活且表达力强的表示空间中生成并演化优化问题。以无约束单目标连续最小化问题为例,这些问题以数学表达式形式表示,旨在显著体现遗传算法(GA)与差分进化(DE)之间的性能差异。实验结果表明,LLM-EBG成功生成的基准问题中,指定目标算法在超过80%的试验中持续优于对比算法。此外,探索性景观分析显示,有利于GA的基准对变量缩放高度敏感,证明该框架能生成反映不同优化算法内在搜索行为的几何特征各异的问题。

原文摘要 · Abstract (English)

Optimization benchmarks play a fundamental role in assessing algorithm performance; however, existing artificial benchmarks often fail to capture the diversity and irregularity of real-world problem structures, while benchmarks derived from real-world problems are costly and difficult to construct. To address these challenges, we propose an evolutionary automatic benchmark generation framework that leverages a large language model (LLM) as a generative operator, termed the LLM-driven evolutionary benchmark generator (LLM-EBG). In this framework, the LLM serves as an evolutionary operator that generates and evolves benchmark problems within a flexible, expressive representation space. As a case study, we generate unconstrained single-objective continuous minimization problems represented as mathematical expressions designed to induce significant performance differences between a genetic algorithm (GA) and differential evolution (DE). Experimental results show that LLM-EBG successfully produces benchmark problems in which the designated target algorithm consistently outperforms the comparative algorithm in more than 80\% of trials. Furthermore, exploratory landscape analysis reveals that benchmarks favoring GA are highly sensitive to variable scaling, demonstrating that the proposed framework can generate problems with distinct geometric characteristics that reflect the intrinsic search behaviors of different optimization algorithms.

优化基准大模型算法评测进化计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。