用视频分析智能体行为,自动设计更有效的强化学习训练任务。
Open-ended Multi-agent Autocurricula via Visual Inspection of Policies with Multi-modal LLMs

- 通过视频语言模型分析智能体行为视频,动态生成训练任务
- 在SMAC上比纯文本或数值评分方法提升训练效率
- 适合研究多智能体强化学习与自动化课程设计的学者
开放式的强化学习课程旨在通过识别能促进技能逐步提升的任务,训练具备通用能力的智能体。设计此类课程的一大挑战在于评估任务难度与智能体当前学习进度的匹配度。以往工作尝试使用标量任务评分或行为文本摘要,本文提出新方法:直接通过录制的智能体运行视频来观察策略行为。我们引入一种简单而有效的方法——视觉策略检视(VIP),利用视频语言模型(VLM)处理视频并生成课程建议。由于视频可自然包含多个可控智能体,我们在星海争霸多智能体挑战(SMAC)上验证了该方法。结果表明,即使使用轻量级开源模型(VideoLLaMa2-7B),VIP生成的课程也优于仅依赖文本或标量评分的基线方法。
原文摘要 · Abstract (English)
Open-ended curricula in Reinforcement Learning (RL) aim to train generally-capable agents by identifying tasks that facilitate learning increasingly complex skills. A major challenge when designing such curricula is assessing task difficulty relative to the agent's current learning progress. While previous work has explored using scalar task scores or textual summaries of the agent's behavior, here we study a different approach: directly inspecting policy behavior via recorded episode videos. We introduce a simple yet effective instantiation of this approach which leverages a Video Language Model (VLM) to both process these videos and provide curriculum recommendations, which we call Visual Inspection of Policies (VIP). Since videos can naturally contain any number of controllable agents, we empirically study VIP on the StarCraft Multi-Agent Challenge (SMAC). We show that even with a lightweight and openly accessible VLM (VideoLLaMa2-7B), VIP can use policy videos to generate more effective curricula than both its text-only ablation and methods that rely on scalar task scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。