arXiv:2602.19532cs.ROcs.SY2026-02被引 6

用贝尔曼值分解解决复杂任务中的安全与目标平衡问题

Bellman Value Decomposition for Task Logic in Safe Optimal Control

  • 将时序逻辑任务的贝尔曼值分解为图结构,通过三类贝尔曼方程连接
  • 在高维非线性系统中实现自动安全与目标兼顾,性能优于基线方法
  • 适合需要多智能体协同、动态约束的机器人控制场景

现实任务涉及目标与安全规范的精细组合。在高维空间中,形式化自动机变得繁琐,稀疏奖励常需大量调参。本文证明:在时序逻辑定义的复杂任务中,贝尔曼值可分解为由已知贝尔曼方程(可达-避障方程、避障方程及新提出的可达-避障-循环方程)连接的值图结构。为此,我们提出VDPPO算法,将分解后的值图嵌入双层神经网络,隐式建模依赖关系。在多种模拟与硬件实验中,针对包含异构团队和非线性动力学的复杂高维任务进行测试,结果表明该方法显著提升性能,能自动平衡安全与活跃性。

原文摘要 · Abstract (English)

Real-world tasks involve nuanced combinations of goal and safety specifications. In high dimensions, the challenge is exacerbated: formal automata become cumbersome, and the combination of sparse rewards tends to require laborious tuning. In this work, we consider the innate structure of the Bellman Value as a means to naturally organize the problem for improved automatic performance. Namely, we prove the Bellman Value for a complex task defined in temporal logic can be decomposed into a graph of Bellman Values, connected by a set of well-known Bellman equations (BEs): the Reach-Avoid BE, the Avoid BE, and a novel type, the Reach-Avoid-Loop BE. To solve the Value and optimal policy, we propose VDPPO, which embeds the decomposed Value graph into a two-layer neural net, bootstrapping the implicit dependencies. We conduct a variety of simulated and hardware experiments to test our method on complex, high-dimensional tasks involving heterogeneous teams and nonlinear dynamics. Ultimately, we find this approach greatly improves performance over existing baselines, balancing safety and liveness automatically.

强化学习安全控制时序逻辑多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。