arXiv:2508.01561cs.AI2025-08NeurIPS被引 12

让AI一次只完成一个子目标,零样本适应复杂时序任务

One Subgoal at a Time: Zero-Shot Generalization to Arbitrary Linear Temporal Logic Requirements in Multi-Task Reinforcement Learning

  • 将时序逻辑任务分解为逐个可达-避障子目标求解
  • 在未见过的时序逻辑规范上实现显著更优的零样本泛化
  • 适合需要安全、灵活执行复杂任务的机器人系统

在强化学习中,泛化到复杂且具有时间延展性的任务目标与安全约束仍是重大挑战。线性时序逻辑(LTL)提供统一形式来定义此类要求,但现有方法难以处理嵌套的长周期任务与安全约束,也无法识别子目标不可满足的情况并切换替代方案。本文提出GenZ-LTL,一种实现任意LTL规范零样本泛化的算法。该方法利用布赫伊自动机结构,将LTL任务分解为一系列可达-避障子目标。不同于当前最优方法依赖子目标序列条件,我们发现通过合理的安全强化学习框架,逐个求解子目标更有效。此外,提出一种新型子目标诱导的观测压缩技术,在合理假设下缓解子目标状态组合的指数级复杂度。实验表明,GenZ-LTL在未见的LTL规范上显著优于现有方法。

原文摘要 · Abstract (English)

Generalizing to complex and temporally extended task objectives and safety constraints remains a critical challenge in reinforcement learning (RL). Linear temporal logic (LTL) offers a unified formalism to specify such requirements, yet existing methods are limited in their abilities to handle nested long-horizon tasks and safety constraints, and cannot identify situations when a subgoal is not satisfiable and an alternative should be sought. In this paper, we introduce GenZ-LTL, a method that enables zero-shot generalization to arbitrary LTL specifications. GenZ-LTL leverages the structure of Büchi automata to decompose an LTL task specification into sequences of reach-avoid subgoals. Contrary to the current state-of-the-art method that conditions on subgoal sequences, we show that it is more effective to achieve zero-shot generalization by solving these reach-avoid problems \textit{one subgoal at a time} through proper safe RL formulations. In addition, we introduce a novel subgoal-induced observation reduction technique that can mitigate the exponential complexity of subgoal-state combinations under realistic assumptions. Empirical results show that GenZ-LTL substantially outperforms existing methods in zero-shot generalization to unseen LTL specifications.

强化学习时序逻辑零样本泛化安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。