arXiv:2503.05042cs.LGcs.AI2025-03被引 6

提出可证明正确的自动机嵌入方法,确保多任务强化学习最优性。

Provably Correct Automata Embeddings for Optimal Automata-Conditioned Reinforcement Learning

  • 基于理论框架设计可验证的自动机嵌入学习方法
  • 证明该设置下多任务策略学习是近似正确可学习的
  • 适合需要可靠时序决策的自动化系统开发者

自动机条件强化学习在运行时通过预训练并冻结自动机嵌入,已展现出学习具备时间扩展目标能力的多任务策略的潜力。然而此前缺乏理论保证。本文构建了自动机条件强化学习问题的理论框架,证明其为可能近似正确可学习(PAC-learnable)。进而提出一种学习可证明正确的自动机嵌入的技术,确保多任务策略学习达到最优。实验结果验证了理论结论的有效性。

原文摘要 · Abstract (English)

Automata-conditioned reinforcement learning (RL) has given promising results for learning multi-task policies capable of performing temporally extended objectives given at runtime, done by pretraining and freezing automata embeddings prior to training the downstream policy. However, no theoretical guarantees were given. This work provides a theoretical framework for the automata-conditioned RL problem and shows that it is probably approximately correct learnable. We then present a technique for learning provably correct automata embeddings, guaranteeing optimal multi-task policy learning. Our experimental evaluation confirms these theoretical results.

强化学习自动机理论保证多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。