arXiv:2503.09620math.OCcs.AI2025-03NAACL被引 7

用编辑大模型的方法,让科学优化更高效可靠。

Exploiting Edited Large Language Models as General Scientific Optimizers

  • 双层优化:模拟器评估+大模型改进,通过编辑更新知识
  • 在7个任务上优于6种大模型,表现稳定且泛化强
  • 适合需要反复试错的科研优化场景,如材料设计、药物发现

大语言模型(LLM)因其丰富的知识和推理能力,被广泛用于科学场景下的数学优化。现有方法多依赖提示工程,以观测反馈作为文本描述,但因大模型对提示高度敏感且易在长提示中迷失,难以有效利用每步优化的反馈,限制了实际应用。为此,我们提出一种概念简单且通用的双层优化方法——通用科学优化器(GSO)。GSO首先使用内层模拟器作为实验平台评估当前解并提供观测反馈;随后,大模型作为具备知识的科学家,基于反馈修正潜在错误生成新解,作为外层优化。最后,通过模型编辑实现模拟器与大模型知识的双向协同更新。大量实验表明,GSO在七项不同任务上,使用六种不同的大模型骨干网络,均持续优于现有最先进方法,验证了其有效性与广泛应用潜力。

原文摘要 · Abstract (English)

Large language models (LLMs) have been widely adopted in mathematical optimization in scientific scenarios for their extensive knowledge and advanced reasoning capabilities. Existing methods mainly focus on utilizing LLMs to solve optimization problems in a prompt-based manner, which takes observational feedback as additional textual descriptions. However, due to LLM's \textbf{high sensitivity to the prompts} and \textbf{tendency to get lost in lengthy prompts}, these methods struggle to effectively utilize the {observational} feedback from each optimization step, which severely hinders the applications for real-world scenarios. To address these challenges, we propose a conceptually simple and general {bi-level} optimization method, namely \textbf{G}eneral \textbf{S}cientific \textbf{O}ptimizers (GSO). Specifically, GSO first utilizes inner-level simulators as experimental platforms to evaluate the current solution and provide observational feedback. Then, LLMs serve as knowledgeable and versatile scientists, generating new solutions by refining potential errors from the feedback as the outer-level optimization. Finally, simulations together with the expert knowledge in LLMs are jointly updated with bi-level interactions via model editing. Extensive experiments show that GSO consistently outperforms existing state-of-the-art methods using \textit{six} different LLM backbones on \textit{seven} different tasks, demonstrating the effectiveness and a wide range of applications.

科学优化大模型双层优化模型编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。