测试大模型在复杂搜索中自我优化的能力,发现越强的模型越会改进但仍未达人类水平。
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces

- 构建新基准,评估模型通过反馈迭代优化解法的能力
- 19个模型实验显示更强模型更善利用反馈,但仍有上限
- 适合研究智能体自适应、认知机制或评测大模型潜力的学者
大型语言模型(LLMs)在推理和工具使用方面表现出色,但其解决复杂问题所依赖的核心认知能力——感知、推理与记忆——是否具备持续自我优化的潜力仍不明确。为此,我们提出OPT-BENCH,一个用于评估大模型在大规模搜索空间中自我改进能力的基准。该基准融合20个机器学习任务与10个经典NP-hard问题,提供严格环境以检验智能体能否通过内在自省而非机械调用工具实现适应性优化。我们进一步提出OPT-Agent框架,模拟人类认知循环:感知-记忆-推理,基于环境反馈迭代改进方案。在涵盖7个模型家族、参数量从3B到235B的19个主流大模型上进行广泛实验,结果表明更强模型更能有效利用反馈信号进行自我提升,但其上限仍由基础能力决定,即使最先进模型仍显著落后于人类专家表现。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning and tool use. However, the fundamental cognitive faculties essential for problem solving, including perception, reasoning, and memory, remain the stable core of intelligence. Unlike memorizing specific patterns, humans succeed in novel environments by applying these intrinsic faculties to adapt and optimize. Yet, whether LLMs possess this essential capacity, namely the ability to continuously refine solutions in response to dynamic environmental feedback, remains underexplored. To address this challenge, we introduce OPT-BENCH, a benchmark for evaluating self-improvement capabilities in large-scale search spaces. By combining 20 machine learning tasks with 10 classic NP-hard problems, OPT-BENCH provides a rigorous setting to assess whether agents can adapt through intrinsic self-reflection rather than rote tool application. We further propose OPT-Agent, a framework that emulates human-like cognitive adaptation. It operates through a general perception, memory, and reasoning loop, iteratively refining solutions based on environmental feedback. Through extensive experiments on 19 LLMs from 7 model families, including reasoning models, general models, and open-source models ranging from 3B to 235B parameters, we demonstrate that stronger models are more effective at leveraging feedback signals for self-improvement. However, this upper-bound adaptability remains fundamentally constrained by the models' base capacity, and even the most advanced LLMs still fall short of human expert performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。