arXiv:2505.03561cs.LGcs.AI2025-05ICML被引 3

提出新型生成流模型,解决连续训练与模仿学习中的损失难题。

Ergodic Generative Flows

  • 基于遍历性设计有限全局变换的生成流,保证通用性且损失可计算。
  • 引入KL-弱流匹配损失,实现无需独立奖励模型的模仿学习。
  • 在2D模拟和真实航天数据上验证有效,适用于强化学习与模仿学习场景。

生成流网络(GFNs)最初用于有向无环图以从非归一化分布中采样。近期研究扩展了生成方法的理论框架,提升了灵活性与应用范围,但在连续设置下训练GFNs及模仿学习(IL)仍面临挑战,包括流匹配损失不可计算、非无环训练测试有限,以及模仿学习需额外奖励模型等问题。本文提出一类名为遍历生成流(EGFs)的生成流家族,以解决上述问题。首先,利用遍历性构建具有有限全局变换(微分同胚)的简单生成流,具备通用性保证且流匹配损失(FM loss)可计算。其次,提出一种结合交叉熵与弱流匹配控制的新损失——KL-弱FM损失,专用于无需独立奖励模型的模仿学习训练。我们在二维模拟任务和来自NASA的球面真实数据集上评估了基于KL-弱FM损失的模仿学习EGFs。此外,还通过目标奖励进行了二维强化学习实验,使用FM损失进行训练。

原文摘要 · Abstract (English)

Generative Flow Networks (GFNs) were initially introduced on directed acyclic graphs to sample from an unnormalized distribution density. Recent works have extended the theoretical framework for generative methods allowing more flexibility and enhancing application range. However, many challenges remain in training GFNs in continuous settings and for imitation learning (IL), including intractability of flow-matching loss, limited tests of non-acyclic training, and the need for a separate reward model in imitation learning. The present work proposes a family of generative flows called Ergodic Generative Flows (EGFs) which are used to address the aforementioned issues. First, we leverage ergodicity to build simple generative flows with finitely many globally defined transformations (diffeomorphisms) with universality guarantees and tractable flow-matching loss (FM loss). Second, we introduce a new loss involving cross-entropy coupled to weak flow-matching control, coined KL-weakFM loss. It is designed for IL training without a separate reward model. We evaluate IL-EGFs on toy 2D tasks and real-world datasets from NASA on the sphere, using the KL-weakFM loss. Additionally, we conduct toy 2D reinforcement learning experiments with a target reward, using the FM loss.

生成模型模仿学习流匹配强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。