从不同目标的示范者中提炼通用行为,提升智能体泛化能力。
Do as the Romans Do: Learning Universal Behaviors from Heterogeneous Agents

- 通过分解奖励函数,分离出跨代理共有的通用行为信号。
- 在合成任务、多智能体游戏和自动驾驶仿真中性能超越基准方法。
- 适合需要高效适应新任务的通用智能体预训练场景。
人类常通过观察他人习得新技能,因为观察行为隐含了环境中的有效行动方式。然而,来自异质群体的示范引入冲突的行为信号,难以判断哪些行为值得模仿。本文提出通用奖励解耦与推断(GRID)方法,从追求不同目标的示范者中提取普遍适用的行为。GRID将每个代理的奖励函数分解为通用奖励(捕捉所有代理共享的行为)和特定奖励(捕捉个体偏好)。仅基于通用奖励训练可形成新型通用智能体预训练范式,使其内化安全性和基础任务能力等通用环境技能,避免标准学习从示范技术中的模式平均偏差。该通用智能体作为下游任务微调的优越先验,可有效适应训练中未见的偏好。在合成基函数分解、多智能体Craftax及连续自动驾驶模拟器Highway-Env上的实验表明,GRID能语义合理地解耦奖励结构,优于标准学习从示范基线,并实现更高效稳定的专精化。
原文摘要 · Abstract (English)
Humans often acquire new skills by observing others, since observed behaviors implicitly reveal how to act in an environment. However, observations drawn from a heterogeneous population introduce conflicting behavioral signals, making it difficult to determine which behaviors are worth imitating. We address this challenge with General Reward Inference and Disentanglement (GRID), a social learning method that extracts universally useful behaviors from a heterogeneous population of demonstrators pursuing different goals. GRID decomposes per-agent reward functions into a general reward, capturing behaviors shared across all agents, and specific rewards, capturing individual preferences and objectives. Training exclusively on the general reward provides a new paradigm of generalist pretraining. It yields a generalist agent that internalizes universal environmental competencies, such as safety and basic task proficiency, without the mode-averaging bias that afflicts standard learning from demonstration techniques. This generalist serves as a superior prior for fine-tuning to downstream tasks, including preferences unseen during training. Experiments across a synthetic basis function decomposition, multi-agent Craftax, and a continuous autonomous driving simulator (Highway-Env) confirm that GRID successfully disentangles reward structure in a semantically meaningful way, outperforms standard learning from demonstration baselines, and enables more efficient and stable specialization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。