arXiv:2604.08779cs.LG2026-04

用自适应模拟实验优化大模型决策策略,提升运营效率。

Adaptive Simulation Experiment for LLM Policy Optimization

论文配图:Adaptive Simulation Experiment for LLM Policy Optimization
图 1 · 摘自论文原文
  • 将大模型视为随机模拟器,通过配对比较寻找最优策略
  • 在两种策略空间下均证明了识别最优策略所需最小数据量
  • 提出LLM-PO算法,比现有方法更高效且具理论保障

大型语言模型(LLMs)在运营管理中具有显著潜力,可提升操作效率。部署需设定影响响应质量、用户体验与运营价值的策略。本文将LLMs视为随机模拟器,提出基于配对比较的自适应模拟实验框架,从有限候选策略中识别最优策略。考虑两种策略空间:无参数假设的非结构化空间,以及由偏好模型生成数据的结构化空间。针对两种情形,分别刻画了以高概率识别最优策略所需的基本数据要求。在非结构化情况下,推导出最优采样比例的闭式表达式,并给出明确的操作解释;在结构化情况下,构建正则化凸规划求解最优比例。进而设计适用于两类空间的自适应实验流程LLM-PO,并证明其在渐近意义上达到最低数据需求的同时,仍能以期望的统计保证识别最优策略。数值实验表明,LLM-PO持续优于基准方法,有效提升LLM性能。

原文摘要 · Abstract (English)

Large language models (LLMs) have significant potential to improve operational efficiency in operations management. Deploying these models requires specifying a policy that governs response quality, shapes user experience, and influences operational value. In this research, we treat LLMs as stochastic simulators and propose a pairwise comparison-based adaptive simulation experiment framework for identifying the optimal policy from a finite set of candidates. We consider two policy spaces: an unstructured space with no parametric assumption, and a structured space in which the data are generated from a preference model. For both settings, we characterize the fundamental data requirements for identifying the optimal policy with high probability. In the unstructured case, we derive a closed-form expression for the optimal sampling proportions, together with a clear operational interpretation. In the structured case, we formulate a regularized convex program to compute the optimal proportions. We then develop an adaptive experimental procedure, termed LLM-PO, for both policy spaces, and prove that it identifies the optimal policy with the desired statistical guarantee while asymptotically attaining the fundamental data requirements. Numerical experiments demonstrate that LLM-PO consistently outperforms benchmark methods and improves LLM performance.

大模型策略模拟实验优化算法决策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。