arXiv:2512.01047cs.AIcs.LG2025-12

自动优化强化学习的逻辑规范,让复杂任务更容易学会。

Automating the Refinement of Reinforcement Learning Specifications

  • 用探索引导策略自动改进粗略的逻辑规范
  • 四类修正方法保持原规范有效性,提升学习成功率
  • 适合需要精准约束的复杂控制任务研究者

逻辑规范已被证明有助于强化学习算法完成复杂任务。然而当任务规范不充分时,智能体可能无法学习到有效策略。本文提出 AutoSpec 框架,通过探索引导策略改进粗粒度的逻辑规范。该框架搜索一种新的规范细化形式,其满足性可保证原始规范成立,同时提供额外指导,使强化学习算法更易学习有效策略。AutoSpec 适用于以 SpectRL 逻辑表达的强化学习任务。利用 SpectRL 规范的组合特性,设计了四种细化操作:修改已有边的规范或引入新边规范。证明了所有四类操作均保持规范正确性——任何满足细化后规范的轨迹也满足原规范。实验表明,使用 AutoSpec 生成的细化规范,显著提升了可解决控制任务的复杂度。

原文摘要 · Abstract (English)

Logical specifications have been shown to help reinforcement learning algorithms in achieving complex tasks. However, when a task is under-specified, agents might fail to learn useful policies. In this work, we explore the possibility of improving coarse-grained logical specifications via an exploration-guided strategy. We propose AutoSpec, a framework that searches for a logical specification refinement whose satisfaction implies satisfaction of the original specification, but which provides additional guidance therefore making it easier for reinforcement learning algorithms to learn useful policies. AutoSpec is applicable to reinforcement learning tasks specified via the SpectRL specification logic. We exploit the compositional nature of specifications written in SpectRL, and design four refinement procedures that modify the abstract graph of the specification by either refining its existing edge specifications or by introducing new edge specifications. We prove that all four procedures maintain specification soundness, i.e. any trajectory satisfying the refined specification also satisfies the original. We then show how AutoSpec can be integrated with existing reinforcement learning algorithms for learning policies from logical specifications. Our experiments demonstrate that AutoSpec yields promising improvements in terms of the complexity of control tasks that can be solved, when refined logical specifications produced by AutoSpec are utilized.

强化学习逻辑规范自动优化SpectRL

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。