arXiv:2511.19885cs.MAcs.LG2025-11AAAI

让游戏智能体读懂抽象战术指令并执行多样风格动作

Complex Instruction Following with Diverse Style Policies in Football Games

  • 用风格参数统一控制多种行为,训练单一策略应对复杂场景
  • 在5v5足球环境中成功理解并执行抽象战术指令
  • 适合需要多风格智能体的复杂协作任务研究

尽管语言控制强化学习(LC-RL)在基础任务和简单指令(如物体操作与导航)中取得进展,但将其扩展至复杂多智能体环境(如足球比赛)中理解和执行高层或抽象指令仍面临重大挑战。为此,我们提出语言控制多样风格策略(LCDSP),一种专为复杂场景设计的新型LC-RL范式。LCDSP包含两个核心组件:多样风格训练(DST)方法与风格解释器(SI)。DST通过调节风格参数(SP)高效训练单一策略,使其能呈现丰富多样的行为模式;SI则可快速准确地将高层语言指令转化为对应的风格参数。在复杂的5v5足球环境中进行大量实验表明,LCDSP能有效理解抽象战术指令,并精确执行期望的多样化行为风格,展现出在复杂真实应用中的潜力。

原文摘要 · Abstract (English)

Despite advancements in language-controlled reinforcement learning (LC-RL) for basic domains and straightforward commands (e.g., object manipulation and navigation), effectively extending LC-RL to comprehend and execute high-level or abstract instructions in complex, multi-agent environments, such as football games, remains a significant challenge. To address this gap, we introduce Language-Controlled Diverse Style Policies (LCDSP), a novel LC-RL paradigm specifically designed for complex scenarios. LCDSP comprises two key components: a Diverse Style Training (DST) method and a Style Interpreter (SI). The DST method efficiently trains a single policy capable of exhibiting a wide range of diverse behaviors by modulating agent actions through style parameters (SP). The SI is designed to accurately and rapidly translate high-level language instructions into these corresponding SP. Through extensive experiments in a complex 5v5 football environment, we demonstrate that LCDSP effectively comprehends abstract tactical instructions and accurately executes the desired diverse behavioral styles, showcasing its potential for complex, real-world applications.

语言控制多智能体足球游戏策略多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。