arXiv:2606.17591cs.AI2026-06中稿 · ICML被引 4

让大模型从真实反馈中学习,通过闭环治理避免过时经验干扰。

Closing the Feedback Loop: From Experience Extraction to Insight Governance in Verbal Reinforcement Learning

  • 构建规则-证据-技能三层结构,用反馈循环管理经验
  • 在金融预测任务中,有治理机制时准确率显著提升
  • 适合需要持续学习真实世界反馈的智能体系统

无需训练的言语强化学习使大模型代理能够从世界反馈(如动态任务结果、市场回报或需求预测)中学习,通过提取经验中的语言规则并注入上下文来更新行为,而无需改变参数。然而,在非平稳环境中,这些代理面临保留与遗忘的困境:保留过时洞察会导致负迁移,舍弃则在条件重现时引发灾难性遗忘。我们识别出四项关键需求——结果驱动评估、持久结构化证据、非单调知识生命周期和组合式治理,并指出现有方法过度投入经验提取,而忽视洞察治理。为此,我们提出一个三层架构——规则、证据、技能——通过反馈驱动的梳理循环闭合治理缺口。规则从世界结果中提炼经验;证据日志追踪每条规则在多轮中的可靠性;技能决定何时应用规则、如何解决冲突以及何时回避。以金融预测为例,世界反馈自然丰富、嘈杂且非平稳,在相同积累经验下,是否具备梳理循环导致性能从低于零样本基线到显著提升准确率和风险调整后收益的截然不同结果。

原文摘要 · Abstract (English)

Training-free verbal reinforcement learning enables LLM agents to learn from world feedback -- objective signals such as dynamic task outcomes, market returns, or demand forecasts -- by extracting verbal rules from experience and injecting them as context, updating the agent's behavior without parameter changes. However, in non-stationary environments these agents face a retention-forgetting dilemma: retaining stale insights causes negative transfer, while discarding them causes catastrophic forgetting when conditions recur. We identify four requirements for navigating this dilemma -- outcome-driven evaluation, persistent structured evidence, non-monotonic knowledge lifecycle, and compositional governance -- and show that existing methods invest heavily in experience extraction while underinvesting in insight governance. We propose a three-layer architecture -- rules, evidence, and skills -- connected by a feedback-driven curation loop that closes the governance gap. Rules capture distilled experience from world outcomes; evidence logs track each rule's reliability across episodes; skills govern which rules to apply, how to resolve conflicts, and when to abstain. On financial forecasting as a case study, where world feedback is naturally abundant, noisy, and non-stationary, we show that the same accumulated experience either degrades performance below the zero-shot baseline or dramatically improves accuracy and risk-adjusted returns, depending on whether the curation loop is present.

强化学习大模型反馈闭环金融预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。