arXiv:2602.09761cs.LGcs.AI2026-02被引 1

让强化学习智能体在视觉环境中理解复杂时序指令并零样本泛化。

Grounding LTL Tasks in Sub-Symbolic RL Environments for Zero-Shot Generalization

  • 用统一经验联合训练策略与符号接地模型,无需预先知道符号对应关系。
  • 在视觉任务中表现接近已知符号接地的基准,显著优于其他无符号先验方法。
  • 适合需要跨任务泛化的自主系统研发人员参考。

本文研究如何在子符号环境(如基于视觉)中训练强化学习智能体,以遵循由线性时序逻辑(LTL)表达的多个时序扩展指令。以往多任务方法通常依赖于观测与公式中符号之间的映射知识,这一假设不现实。本文通过共享经验,联合训练多任务策略与符号接地模型,后者仅从原始观测和稀疏奖励中,利用神经奖励机以半监督方式学习。在基于视觉的环境中实验表明,该方法性能接近使用真实符号接地的基线,且显著优于唯一其他无需符号先验的多任务学习方法。

原文摘要 · Abstract (English)

In this work we address the problem of training a Reinforcement Learning agent to follow multiple temporally-extended instructions expressed in Linear Temporal Logic in sub-symbolic environments. Previous multi-task work has mostly relied on knowledge of the mapping between raw observations and symbols appearing in the formulae. We drop this unrealistic assumption by jointly training a multi-task policy and a symbol grounder with the same experience. The symbol grounder is trained only from raw observations and sparse rewards via Neural Reward Machines in a semi-supervised fashion. Experiments on vision-based environments show that our method achieves performance comparable to using the true symbol grounding and significantly outperforms the only other previous method for multi-task learning that does not assume knowledge of the true symbol grounding.

强化学习时序逻辑零样本泛化符号接地

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。