arXiv:2609.03177cs.LG2026-09

前沿大模型可高效优化离散任务,但连续优化表现不稳定。

Frontier LLMs are effective batch optimizers: Assessing reasoning models in continuous and discrete settings

  • 用大模型做批量优化,零样本直接上手。
  • 在数值测试中表现接近传统方法,但易出错。
  • 在语义丰富的离散任务中优势明显,适合相似场景。

前沿大语言模型(LLMs)因大规模预训练,具备在多种优化场景中导航的能力,因而成为极具吸引力的优化先验。然而,当前推理型大模型在批量优化任务中的表现仍缺乏系统评估。本文研究了最新一代前沿大模型在连续与离散优化设置下的表现。结果表明,尽管大模型在数值测试函数上作为零样本批量优化器具有竞争力,但其性能相比经典非大模型优化方法更为脆弱。然而,在语义丰富的场景中,大模型先验表现出显著优势,说明其在结构与预训练数据相似的离散空间中,具有极强的优化能力。

原文摘要 · Abstract (English)

Frontier large language models (LLMs) have become attractive priors for optimization due to their large-scale pretraining that enables them to navigate a variety of optimization settings. However, the effectiveness of modern reasoning LLMs in batch optimization settings remains underexplored. Here we investigate the performance of the current generation of frontier LLMs as batch optimizers in both continuous and discrete settings. We find that while LLMs are competitive zero-shot batch optimizers for numerical test functions, their performance is brittle compared to classical non-LLM optimization approaches. However, LLM priors are significantly better in semantically rich settings, indicating that their batch optimization behavior is highly effective when navigating and reasoning over the discrete spaces most similar in structure to their pretraining data.

大模型优化离散优化推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。