arXiv:2503.05818cs.LGcs.RO2025-03被引 5

用逻辑优先级解决多目标强化学习中的意图落差问题。

Closing the Intent-to-Behavior Gap via Fulfillment Priority Logic

  • 通过逻辑公式表达多目标优先级,替代传统线性奖励组合。
  • 在连续控制任务中实现比SAC高500%的样本效率提升。
  • 首次为连续控制设计非线性效用标量化方法,适合复杂机器人任务。

强化学习策略设计面临的核心挑战是将行为意图转化为合理的奖励函数。这种挑战源于意图需同时达成多个相互竞争的目标,传统依赖人工设计的线性奖励组合方法效果脆弱。例如,在机器人任务中,性能最大化与能耗最小化存在天然冲突,简单线性组合难以应对。本文提出基于目标实现(Objective Fulfillment)的新范式,并构建满足此范式的实现优先级逻辑(Fulfillment Priority Logic, FPL)。FPL允许从业者以逻辑公式形式表达其多目标意图与优先级。我们提出的平衡策略梯度算法(Balanced Policy Gradient)利用FPL规范,相较Soft Actor Critic在连续控制任务中实现最高500%的样本效率提升。本工作首次实现了针对连续控制问题的非线性效用标量化设计。

原文摘要 · Abstract (English)

Practitioners designing reinforcement learning policies face a fundamental challenge: translating intended behavioral objectives into representative reward functions. This challenge stems from behavioral intent requiring simultaneous achievement of multiple competing objectives, typically addressed through labor-intensive linear reward composition that yields brittle results. Consider the ubiquitous robotics scenario where performance maximization directly conflicts with energy conservation. Such competitive dynamics are resistant to simple linear reward combinations. In this paper, we present the concept of objective fulfillment upon which we build Fulfillment Priority Logic (FPL). FPL allows practitioners to define logical formula representing their intentions and priorities within multi-objective reinforcement learning. Our novel Balanced Policy Gradient algorithm leverages FPL specifications to achieve up to 500\% better sample efficiency compared to Soft Actor Critic. Notably, this work constitutes the first implementation of non-linear utility scalarization design, specifically for continuous control problems.

强化学习多目标连续控制非线性标量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。