arXiv:2507.10741cs.LGcs.AI2025-07NeurIPS被引 4

用少量数据让智能体从语言指令中学会复杂行为

Ground-Compose-Reinforce: Grounding Language in Agentic Behaviours using Limited Data

  • 通过可组合的奖励机框架,从高阶任务描述直接训练智能体
  • 仅用350条标注轨迹就实现复杂行为生成,非组合方法失败
  • 适合少样本场景下构建能理解语言的交互式智能体

将语言与感知和行动进行对齐,是构建能够通过语言与人类或其他智能体交互的场景化智能体的关键挑战。以往方法需手动设计语言对齐机制或依赖大规模数据集。本文提出一种端到端的神经符号框架 Ground-Compose-Reinforce,可直接从高层任务规范(如奖励机)训练强化学习智能体,无需人工设计奖励函数或领域特定先验,也无需海量数据。这些任务规范以自动化的状态机形式表示,部分可由自然语言自动生成。关键在于,我们证明了通过利用组合性,即使在有限数据条件下也能实现奖励机的对齐。在自定义的 Meta-World 域中,仅使用350条标注预训练轨迹,该框架便能忠实激发复杂行为——包括预训练中从未出现过的动作序列,而无组合性的方法则无法达成。

原文摘要 · Abstract (English)

Grounding language in perception and action is a key challenge when building situated agents that can interact with humans, or other agents, via language. In the past, addressing this challenge has required manually designing the language grounding or curating massive datasets that associate language with the environment. We propose Ground-Compose-Reinforce, an end-to-end, neurosymbolic framework for training RL agents directly from high-level task specifications--without manually designed reward functions or other domain-specific oracles, and without massive datasets. These task specifications take the form of Reward Machines, automata-based representations that capture high-level task structure and are in some cases autoformalizable from natural language. Critically, we show that Reward Machines can be grounded using limited data by exploiting compositionality. Experiments in a custom Meta-World domain with only 350 labelled pretraining trajectories show that our framework faithfully elicits complex behaviours from high-level specifications--including behaviours that never appear in pretraining--while non-compositional approaches fail.

智能体语言对齐少样本强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。