arXiv:2504.00907cs.AI2025-04被引 25

让机器人通过主动提问澄清模糊指令,提升家务任务执行效率。

Grounding Multimodal LLMs to Embodied Agents that Ask for Help with Reinforcement Learning

  • 用强化学习微调多模态大模型,使其能根据视觉语言信息决策提问
  • 在新场景和任务上表现优于基线10.4%-16.5%,无需人工示范
  • 首次实现基于LLM生成奖励的在线强化学习,支持机器人自主求助

在家庭环境中运行的具身智能体必须理解模糊且不完整的用户指令。一个高效的家用机器人应能识别歧义,并提出相关的问题以准确推断用户意图,从而更有效地完成任务。为此,我们提出了「Ask-to-Act」任务,要求智能体在家庭环境中执行单目标或多重目标重排任务,使用不完整指令,同时需在部分可观测条件下,策略性地提出最少但最相关的澄清问题。为应对这一挑战,我们提出一种新方法:利用在线强化学习(RL)对多模态大语言模型(MLLM)进行微调,构建视觉-语言-动作(VLA)策略,采用由大语言模型生成的奖励信号。该方法无需大规模人工示范或人工设计奖励函数。我们在该任务上对比了包括GPT-4o在内的强零样本基线及监督微调的MLLM。结果表明,我们的强化学习微调的MLLM显著优于所有基线(提升10.4%-16.5%),并在新场景与任务中表现出良好泛化能力。据我们所知,这是首次展示通过大语言模型生成奖励并结合在线强化学习,将多模态大模型适配为可行动、会求助的具身智能体。

原文摘要 · Abstract (English)

Embodied agents operating in household environments must interpret ambiguous and under-specified human instructions. A capable household robot should recognize ambiguity and ask relevant clarification questions to infer the user intent accurately, leading to more effective task execution. To study this problem, we introduce the Ask-to-Act task, where an embodied agent is tasked with a single or multi-object rearrangement task using an under-specified instruction in a home environment. The agent must strategically ask minimal, yet relevant, clarification questions to resolve ambiguity while navigating under partial observability. To address this challenge, we propose a novel approach that fine-tunes multi-modal large language models (MLLMs) as vision-language-action (VLA) policies using online reinforcement learning (RL) with LLM-generated rewards. Our method eliminates the need for large-scale human demonstrations or manually engineered rewards for training such agents. We benchmark against strong zero-shot baselines including GPT-4o as well as supervised fine-tuned MLLMs on our task. Our results show that our RL-finetuned MLLM outperforms all baselines by a significant margin (10.4-16.5%), generalizing well to novel scenes and tasks. To the best of our knowledge, this is the first demonstration of adapting MLLMs as VLA agents that can act and ask for help using LLM-generated rewards with online RL.

具身智能多模态大模型强化学习主动提问

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。