测试机器人抓取策略在分布外任务上的归纳能力,发现现有模型看似泛化实则失效。
Inductive Generalization for Robotic Manipulation

- 设计新评估协议,用渐进式分布外任务测试策略的归纳能力
- 现有顶尖视觉语言动作模型在新测试中表现不佳,暴露泛化盲区
- 适合关注机器人真实泛化性能的研究者与开发者
理解视觉运动策略的泛化能力对发展强大机器人代理至关重要。可泛化的模型能学习跨领域迁移的结构。然而,实践中视觉运动策略通常通过已知分布内的插值测试,使用非结构化的域偏移(如光照变化、杂乱环境、多样物体)。我们主张,为衡量泛化能力,应测试策略在逐步增强的分布外任务变体上的归纳能力。这被称为归纳泛化,借鉴了轴向评估揭示语言模型内在泛化限制的方法(如序列长度、计数)。本文提供了一个可复用且形式化的评估协议,用于测量任何抓取策略的归纳泛化能力,并建立基线表明现有范式无法通过此测试;例如,当前最先进的视觉-语言-动作模型在新测试中失败。这些结果揭示了一类与数据和模型规模扩展无关但对实现通用机器人至关重要的学习挑战。
原文摘要 · Abstract (English)
Understanding the generalization capabilities of visuomotor policies is essential in the development of capable robotic agents. Generalizable models learn structures that transfer across domains. However, in practice, visuomotor policies test performance by interpolation on known distributions using unstructured domain shifts (e.g. lighting, clutter, diverse objects). We argue that to measure generalization capabilities we must instead test the inductive capacity of policies on progressively harder, out-of-distribution task variants. We call this inductive generalization, drawing directly on how axis-based evaluation has revealed inherent generalization limitations in language models (e.g. sequence length, counting) arXiv:2502.00197 . We provide a reusable and formal evaluation protocol for measuring inductive generalization in any manipulation policy, and establish baselines showing that existing paradigms fail this test; e.g. SoTA Vision-Language-Action models and find that policies that appear to generalize to prior domain shifts (distractors, etc) fail inductive generalization tests. These results expose a class of learning challenges orthogonal to those addressed by data and model scaling in robot learning, yet are imperative to solve in order to realize general purpose robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。