arXiv:2606.04360cs.CLcs.LG2026-06被引 3

用智能体框架提升大模型符号回归效率,仅用40%数据达更好效果

Deliberate Evolution: Agentic Reasoning for Sample-Efficient Symbolic Regression with LLMs

论文配图:Deliberate Evolution: Agentic Reasoning for Sample-Efficient Symbolic Regression with LLMs
图 1 · 摘自论文原文
  • 分离生成与搜索控制,用智能体分步指导演化方向
  • 在多个科学领域实验中,仅用40%样本即超越现有基线
  • 适合需要高效建模的科研人员和自动化发现场景

符号回归(SR)从数据中发现紧凑的数学表达式,但现有基于大模型的进化方法因依赖均方误差等单一评分,仍存在样本效率低的问题。我们发现核心瓶颈在于:候选表达式生成与搜索引导混杂,导致大模型需单凭一个分数同时完成演化方向判断、错误诊断和经验复用。为此,提出刻意演化(Deliberate Evolution, DE)框架,将符号生成与搜索控制解耦。DE通过自适应算子指引搜索方向,利用分析工具进行结构诊断,并借助反思记忆积累轨迹级经验。在LLM-SRBench上的实验表明,DE在多种科学领域持续优于代表性大模型符号回归基线,且仅使用标准样本预算的40%。

原文摘要 · Abstract (English)

Symbolic regression (SR) discovers compact mathematical expressions from data, yet recent LLM-based evolutionary methods remain sample-inefficient because they rely mainly on scalar feedback such as MSE. We identify a core limitation: existing methods conflate candidate proposal with search guidance, requiring the LLM to infer how to evolve an expression, diagnose its errors, and reuse past experience from a single score. To address this, we propose Deliberate Evolution (DE), an agentic framework that decouples symbolic generation from search control. DE guides LLM proposals with adaptive operators for search direction, analytical tools for structural diagnosis, and reflective memory for trajectory-level experience. Experiments on LLM-SRBench show that DE consistently outperforms representative LLM-based SR baselines across diverse scientific domains while using only 40% of the standard sample budget.

符号回归大模型智能体高效建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。