arXiv:2411.19787cs.LGcs.AI2024-11中稿 · TMLR 2025被引 1

通过跨模态辅助目标提升指令感知,实现高效多模态强化学习

CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives

  • 引入指令追踪机制与视频文本检索启发的辅助损失
  • 在多模态任务中显著提升样本效率与系统泛化能力
  • 适合需要精准理解语言指令的机器人控制场景

将指令与环境进行语义对齐是解决语言引导目标达成强化学习问题的关键步骤。在自动化强化学习中,提升模型在不同任务和环境中的泛化能力是一大挑战。在目标达成场景中,智能体需在环境上下文中理解指令的不同组成部分,才能成功完成整体任务。本文提出 CAREL(跨模态辅助强化学习)框架,采用受视频-文本检索研究启发的辅助损失函数,并引入一种名为指令追踪的新方法,可自动跟踪环境中的进展。实验结果表明,该框架在多模态强化学习问题中展现出更优的样本效率和系统泛化能力。代码已公开。

原文摘要 · Abstract (English)

Grounding the instruction in the environment is a key step in solving language-guided goal-reaching reinforcement learning problems. In automated reinforcement learning, a key concern is to enhance the model's ability to generalize across various tasks and environments. In goal-reaching scenarios, the agent must comprehend the different parts of the instructions within the environmental context in order to complete the overall task successfully. In this work, we propose CAREL (Cross-modal Auxiliary REinforcement Learning) as a new framework to solve this problem using auxiliary loss functions inspired by video-text retrieval literature and a novel method called instruction tracking, which automatically keeps track of progress in an environment. The results of our experiments suggest superior sample efficiency and systematic generalization for this framework in multi-modal reinforcement learning problems. Our code base is available here.

强化学习指令理解跨模态机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。