arXiv:2503.22886cs.LGcs.RO2025-03被引 5

用任务令牌让通用行为模型灵活适配具体任务。

Task Tokens: A Flexible Approach to Adapting Behavior Foundation Models

  • 通过强化学习训练新编码器,将观测映射为任务令牌输入模型。
  • 在多种任务中提升表现,包括分布外场景且保持泛化能力。
  • 适合需要快速适配的机器人控制研究者使用。

近期模仿学习进展催生了基于Transformer的行为基础模型(BFM),使类人机器人实现多模态、自然的控制。尽管能零样本生成稳健行为,但针对特定任务常需精细提示工程,可能导致次优结果。本文提出“任务令牌”方法,有效定制BFM以适应具体任务,同时保留其灵活性。该方法利用BFM的Transformer架构,通过强化学习训练一个新任务编码器,冻结原模型不变。用户可引入先验知识,平衡奖励设计与提示工程。通过训练编码器将观测映射为任务令牌并作为额外输入,引导性能提升,同时保持模型多样化的控制特性。我们在多种任务中验证了任务令牌的有效性,包括分布外场景,并展示了其与其他提示方式的兼容性。结果表明,任务令牌为适配BFM至具体控制任务提供了有前景的方案,同时维持其泛化能力。

原文摘要 · Abstract (English)

Recent advancements in imitation learning have led to transformer-based behavior foundation models (BFMs) that enable multi-modal, human-like control for humanoid agents. While excelling at zero-shot generation of robust behaviors, BFMs often require meticulous prompt engineering for specific tasks, potentially yielding suboptimal results. We introduce "Task Tokens", a method to effectively tailor BFMs to specific tasks while preserving their flexibility. Our approach leverages the transformer architecture of BFMs to learn a new task-specific encoder through reinforcement learning, keeping the original BFM frozen. This allows incorporation of user-defined priors, balancing reward design and prompt engineering. By training a task encoder to map observations to tokens, used as additional BFM inputs, we guide performance improvement while maintaining the model's diverse control characteristics. We demonstrate Task Tokens' efficacy across various tasks, including out-of-distribution scenarios, and show their compatibility with other prompting modalities. Our results suggest that Task Tokens offer a promising approach for adapting BFMs to specific control tasks while retaining their generalization capabilities.

行为模型机器人控制强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。