在无专家动作标签下,用区域诊断反馈修复定价策略,实现接近基准的收益表现。
When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions
- 仅凭区域级诊断反馈,通过多轮大模型编辑修复策略
- 收益达108.47,接近基准108.75,且策略分布更贴近基准
- 适用于缺乏细粒度标注的智能系统自修复场景
代理型AI系统常用于修正决策策略,但当每状态的专家动作标签不可得时,评估修正效果极具挑战。本文在酒店定价模拟器中研究此问题:策略编辑器仅能获取区域级诊断反馈,即其价格分布与基准策略在时间、库存和市场区域上的差异汇总。编辑器无法观测基准动作、源码、奖励数值或保留结果,只能对目标动作表进行受限修改。在5,000个保留回合中,多重启大模型编辑器实现RevPAR 108.47(95% CI 107.61–109.34),接近基准策略的108.75(107.81–109.68),配对差距(LLM减基准)为-0.276,95% CI [-0.692, 0.146]。廉价诊断投影已恢复大部分收益(107.90),因此大模型编辑器的优势不仅在于收益提升,还在于将策略组成距离从1.153降至0.609。这是最强的非基准修复结果。该成效非单纯重启搜索所致:无语义的提议者最多2,500次评估仍落后8.77–14.57 RevPAR。也非提示格式所致:打乱诊断信息后收益骤降至94.30。树状编辑器虽具更强整体对齐(0.214 vs 0.266)与参考状态距离(D1 0.328 vs 1.197),但收益降至98.91。结果表明,代理策略修复应以诊断反馈能否形成可靠闭环结果为评价标准,而非单一行为距离。
原文摘要 · Abstract (English)
Agentic AI systems are increasingly used to edit, refine, and repair decision policies, but evaluating these edits is difficult when per-state expert action labels are unavailable. We study this problem in a hotel-pricing simulator where an agentic policy editor receives only region-level diagnostic feedback: summaries of how its price distribution differs from a benchmark policy across time, inventory, and market regions. The editor cannot observe benchmark actions, benchmark source code, reward numbers, or held-out outcomes, and may only propose constrained edits to a target-action table. On 5,000 held-out episodes, a multi-restart LLM editor reaches RevPAR 108.47 (95% CI 107.61 - 109.34), close to the benchmark policy's 108.75 (107.81 - 109.68), with paired gap (LLM minus benchmark) -0.276 and 95% CI [-0.692, 0.146]. A cheap diagnostic projection already recovers much of the revenue (107.90), so the LLM editor's distinctive gain is not raw revenue lift alone: it also reduces episode composition distance from 1.153 to 0.609. This is the strongest non-benchmark repair result. This profile is not explained by restart search alone: non-semantic proposers with up to 2,500 evaluations fall 8.77 - 14.57 RevPAR points short. Nor is it explained by plausible prompt format: a shuffled-diagnostic control breaks region-error correspondence and falls to RevPAR 94.30. The match is genuine but partial. A tree editor achieves stronger pooled alignment, 0.214 versus 0.266, and stronger reference-state D1, 0.328 versus 1.197, yet revenue falls to 98.91. These results show that agentic policy repair should be evaluated by whether diagnostic feedback becomes reliable closed-loop outcome, not by a single behavioral distance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。