arXiv:2602.20728cs.AI2026-02被引 2

用AI反馈实现多目标交通控制的平衡,无需手动调奖赏。

Balancing Multiple Objectives in Urban Traffic Control with Reinforcement Learning from AI Feedback

  • 通过AI生成偏好标签,自动学习多目标权衡策略。
  • 在不人工设计奖励的情况下,实现不同用户优先级下的平衡控制。
  • 适合需兼顾效率与公平的智能交通系统开发者。

奖励设计是现实世界强化学习部署的核心挑战,尤其在多目标场景中。基于人类偏好的强化学习提供了一种有前景的替代方案,通过比较行为结果来学习。近期,从AI反馈中学习(RLAIF)证明大型语言模型可规模化生成偏好标签,减少对人工标注的依赖。然而,现有RLAIF研究多集中于单目标任务,尚未解决多目标系统中冲突目标间权衡的难题。此类系统难以明确设定权衡标准,政策可能退化为仅优化主导目标。本文探索将RLAIF范式扩展至多目标自适应系统,表明多目标RLAIF能生成反映不同用户优先级的平衡策略,避免繁琐的奖励工程。我们认为,将RLAIF融入多目标强化学习,为具有内在冲突目标的领域提供了可扩展的用户对齐策略学习路径。

原文摘要 · Abstract (English)

Reward design has been one of the central challenges for real world reinforcement learning (RL) deployment, especially in settings with multiple objectives. Preference-based RL offers an appealing alternative by learning from human preferences over pairs of behavioural outcomes. More recently, RL from AI feedback (RLAIF) has demonstrated that large language models (LLMs) can generate preference labels at scale, mitigating the reliance on human annotators. However, existing RLAIF work typically focuses only on single-objective tasks, leaving the open question of how RLAIF handles systems that involve multiple objectives. In such systems trade-offs among conflicting objectives are difficult to specify, and policies risk collapsing into optimizing for a dominant goal. In this paper, we explore the extension of the RLAIF paradigm to multi-objective self-adaptive systems. We show that multi-objective RLAIF can produce policies that yield balanced trade-offs reflecting different user priorities without laborious reward engineering. We argue that integrating RLAIF into multi-objective RL offers a scalable path toward user-aligned policy learning in domains with inherently conflicting objectives.

强化学习交通控制多目标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。