提出新基准与方法,提升机器人跨任务零样本泛化能力。
Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
- 用上下文演示引导大模型预测未见任务的动作序列。
- 在23个未见任务上,新方法显著超越现有模型性能。
- 适合研究通用机器人操作与跨任务泛化的学者。
现有视觉-语言-动作(VLA)模型在未见任务上的跨任务泛化能力尚未充分探索。为此,我们引入AGNOSTOS,一个包含23个与训练分布不同的未见操作任务的仿真评估基准,设置两级泛化难度以测试鲁棒性。系统评估发现,尽管当前VLA模型在多样数据集上训练,仍难以有效泛化至这些未见任务。为此,我们提出跨任务上下文操作(X-ICM)方法,利用大语言模型(LLMs)基于已见任务的上下文示范,预测未见任务的动作序列;同时引入动态引导样本选择策略,通过捕捉跨任务动力学特征筛选相关示范。在AGNOSTOS上,X-ICM显著提升跨任务零样本泛化性能,优于领先VLA模型。我们认为AGNOSTOS与X-ICM将推动通用机器人操作的发展。
原文摘要 · Abstract (English)
The generalization capabilities of vision-language-action (VLA) models to unseen tasks are crucial to achieving general-purpose robotic manipulation in open-world settings. However, the cross-task generalization capabilities of existing VLA models remain significantly underexplored. To address this gap, we introduce AGNOSTOS, a novel simulation benchmark designed to rigorously evaluate cross-task zero-shot generalization in manipulation. AGNOSTOS comprises 23 unseen manipulation tasks for testing, distinct from common training task distributions, and incorporates two levels of generalization difficulty to assess robustness. Our systematic evaluation reveals that current VLA models, despite being trained on diverse datasets, struggle to generalize effectively to these unseen tasks. To overcome this limitation, we propose Cross-Task In-Context Manipulation (X-ICM), a method that conditions large language models (LLMs) on in-context demonstrations from seen tasks to predict action sequences for unseen tasks. Additionally, we introduce a dynamics-guided sample selection strategy that identifies relevant demonstrations by capturing cross-task dynamics. On AGNOSTOS, X-ICM significantly improves cross-task zero-shot generalization performance over leading VLAs. We believe AGNOSTOS and X-ICM will serve as valuable tools for advancing general-purpose robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。