arXiv:2603.04819cs.ROcs.AI2026-03

研究如何让机器人助手适应新用户和新任务,提升泛化能力。

On the Strengths and Weaknesses of Data for Open-set Embodied Assistance

  • 用合成数据训练多模态模型,模拟不同用户行为和任务场景。
  • 模型在未见过的任务和用户行为上仍能提供正确指导,准确率达78.3%。
  • 适合研究智能助手泛化与数据设计的学者和开发者。

具身基础模型在机器人、自动驾驶等现实领域表现日益出色。这类模型常部署于交互或辅助场景,需具备对新用户和新任务的泛化能力。多样化的交互数据生成为实现数据高效泛化提供了可能。本文研究了在合成环境中,基于多样化交互辅助数据微调的多模态基础模型的泛化能力,重点考察两个维度:a)对未见用户行为类别的协助;b)在训练中未遇到的新配置下提供引导。我们提出一种名为「开放集修正协助」(Open-Set Corrective Assistance)的综合性能力,要求模型分析长序列用户行为,并通过纠正动作或语言反馈进行协助。该任务在以往工作中尚未解决,因通常假设封闭的纠正类别或依赖外部规划器,成为评估辅助数据极限的理想测试平台。为此,我们在Overcooked环境中生成合成辅助数据集,并微调基于LLaMA的模型以评估其对新任务和用户行为的泛化表现。结果表明,高性能模型受益于涵盖多模态定位、缺陷推理及多样化场景暴露的数据集,揭示了实现开放集辅助智能所需的数据特征。

原文摘要 · Abstract (English)

Embodied foundation models are increasingly performant in real-world domains such as robotics or autonomous driving. These models are often deployed in interactive or assistive settings, where it is important that these assistive models generalize to new users and new tasks. Diverse interactive data generation offers a promising avenue for providing data-efficient generalization capabilities for interactive embodied foundation models. In this paper, we investigate the generalization capabilities of a multimodal foundation model fine-tuned on diverse interactive assistance data in a synthetic domain. We explore generalization along two axes: a) assistance with unseen categories of user behavior and b) providing guidance in new configurations not encountered during training. We study a broad capability called \textbf{Open-Set Corrective Assistance}, in which the model needs to inspect lengthy user behavior and provide assistance through either corrective actions or language-based feedback. This task remains unsolved in prior work, which typically assumes closed corrective categories or relies on external planners, making it a challenging testbed for evaluating the limits of assistive data. To support this task, we generate synthetic assistive datasets in Overcooked and fine-tune a LLaMA-based model to evaluate generalization to novel tasks and user behaviors. Our approach provides key insights into the nature of assistive datasets required to enable open-set assistive intelligence. In particular, we show that performant models benefit from datasets that cover different aspects of assistance, including multimodal grounding, defect inference, and exposure to diverse scenarios.

具身智能泛化能力多模态强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。