arXiv:2602.08054cs.LGcs.AI2026-02被引 2

用图示引导流匹配,实现离线强化学习的安全与性能协同优化。

Epigraph-Guided Flow Matching for Safe and Performant Offline Reinforcement Learning

  • 基于状态约束最优控制建模,通过上图值函数联合优化安全与性能。
  • 在Safety-Gymnasium上实现近零安全违规,收益媲美先进方法。
  • 适合需高安全性与稳定性的机器人控制、自动驾驶等场景。

离线强化学习为训练自主系统提供了无在线探索风险的范式,尤其适用于安全关键领域。然而,在固定数据集上同时实现强安全性和高性能仍具挑战。现有安全离线强化学习方法常依赖软约束,允许违规、过度保守或难以平衡安全、奖励优化与数据分布一致性。为此,我们提出上图引导流匹配(EpiFlow),将安全离线强化学习建模为状态约束最优控制问题,以协同优化安全与性能。我们学习一个由上图重构导出的可行性值函数,避免了以往方法中目标解耦或事后过滤的弊端。通过该值函数重加权行为分布,并利用流匹配拟合生成策略,实现高效且分布一致的采样。在多种安全关键任务(包括Safety-Gymnasium基准)上,EpiFlow在获得竞争力回报的同时,实现了近乎零的实测安全违规,验证了上图引导策略合成的有效性。

原文摘要 · Abstract (English)

Offline reinforcement learning (RL) provides a compelling paradigm for training autonomous systems without the risks of online exploration, particularly in safety-critical domains. However, jointly achieving strong safety and performance from fixed datasets remains challenging. Existing safe offline RL methods often rely on soft constraints that allow violations, introduce excessive conservatism, or struggle to balance safety, reward optimization, and adherence to the data distribution. To address this, we propose Epigraph-Guided Flow Matching (EpiFlow), a framework that formulates safe offline RL as a state-constrained optimal control problem to co-optimize safety and performance. We learn a feasibility value function derived from an epigraph reformulation of the optimal control problem, thereby avoiding the decoupled objectives or post-hoc filtering common in prior work. Policies are synthesized by reweighting the behavior distribution based on this epigraph value function and fitting a generative policy via flow matching, enabling efficient, distribution-consistent sampling. Across various safety-critical tasks, including Safety-Gymnasium benchmarks, EpiFlow achieves competitive returns with near-zero empirical safety violations, demonstrating the effectiveness of epigraph-guided policy synthesis.

强化学习安全控制离线学习流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。