用视觉语言模型让多智能体学会符合人类常识,自动调整策略。
Vision-Based Generic Potential Function for Policy Alignment in Multi-Agent Reinforcement Learning
- 底层用VLM作通用势函数,通过语义理解引导策略对齐常识
- 顶层动态选择合适势函数,适应长时任务中的不确定性变化
- 理论证明不改变最优策略,适合复杂协作场景
在复杂且长时程的多智能体强化学习任务中,使策略对齐人类常识是一个难题,主要源于常识难以建模为奖励信号。现有方法依赖专家设计规则奖励,成本高且缺乏高层语义理解。为此,我们提出一种分层视觉奖励塑造方法:底层使用视觉语言模型(VLM)作为通用势函数,通过内在语义理解引导策略对齐人类常识;顶层基于视觉大语言模型(vLLM)设计自适应技能选择模块,利用指令、视频回放和训练记录从预设池中动态选择合适的势函数以应对长时任务中的不确定性。该方法在理论层面保证最优策略不变。在Google Research Football环境的大量实验表明,该方法不仅显著提升胜率,还有效实现策略与人类常识对齐。
原文摘要 · Abstract (English)
Guiding the policy of multi-agent reinforcement learning to align with human common sense is a difficult problem, largely due to the complexity of modeling common sense as a reward, especially in complex and long-horizon multi-agent tasks. Recent works have shown the effectiveness of reward shaping, such as potential-based rewards, to enhance policy alignment. The existing works, however, primarily rely on experts to design rule-based rewards, which are often labor-intensive and lack a high-level semantic understanding of common sense. To solve this problem, we propose a hierarchical vision-based reward shaping method. At the bottom layer, a visual-language model (VLM) serves as a generic potential function, guiding the policy to align with human common sense through its intrinsic semantic understanding. To help the policy adapts to uncertainty and changes in long-horizon tasks, the top layer features an adaptive skill selection module based on a visual large language model (vLLM). The module uses instructions, video replays, and training records to dynamically select suitable potential function from a pre-designed pool. Besides, our method is theoretically proven to preserve the optimal policy. Extensive experiments conducted in the Google Research Football environment demonstrate that our method not only achieves a higher win rate but also effectively aligns the policy with human common sense.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。