arXiv:2606.00840cs.AI2026-06

用逻辑验证框架评估强化学习在新任务上的泛化能力。

Certificate-Guided Evaluation of Reinforcement Learning Generalization

论文配图:Certificate-Guided Evaluation of Reinforcement Learning Generalization
图 1 · 摘自论文原文
  • 构建结构相似的可达-规避任务族,用于测试泛化性能。
  • 引入神经证书函数,通过轨迹验证发现算法泛化缺陷。
  • 证书违规率越低,成功解决的任务越多,适合作为通用评估标准。

本文提出一种基于逻辑的框架,用于评估强化学习(RL)算法在未见任务上的泛化能力。该框架定义了一类具有结构相似性的归纳可达-规避任务,能够衡量算法在复杂连续环境中的泛化表现。我们引入一种神经证书函数,通过强制执行关键条件来验证RL算法生成的轨迹,从而作为检验泛化能力的试金石。我们在多个先进可泛化RL算法上进行了实证评估,结果表明证书函数违规率越低,成功解决的测试任务数量越多,证明了该框架在区分和评估不同算法泛化能力方面的有效性。本工作为强化学习泛化提供了一个系统性评测方法。

原文摘要 · Abstract (English)

This work presents a logic-driven framework to evaluate the performance of reinforcement learning (RL) algorithms in their ability to generalize to unseen tasks. Our framework defines a family of inductive reach-avoid tasks, characterized by structural similarities in task dynamics, enabling evaluation of generalization capabilities. We introduce a neural certificate function that validates trajectories generated by RL algorithms by enforcing key conditions, thereby serving as a litmus test for RL generalization. We empirically demonstrate our method's capability in certifying generalization for several state-of-the-art generalizable RL algorithms on challenging continuous environments. Our results show that a lower percentage of certificate function violations correlates with a higher number of test tasks successfully solved, highlighting the effectiveness of our framework in evaluating and distinguishing generalization capabilities of RL algorithms. This work provides a principled approach for benchmarking RL generalization.

强化学习泛化评估证书验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。