发现机器人通用策略因数据多样性不足而依赖表面特征,影响泛化能力。
Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
- 分析指出数据集内部多样性差与跨子集分布差异是导致捷径学习主因
- 实验证明在仿真和真实环境中的策略泛化性能显著提升
- 提出数据增强方法,适合资源受限下优化现有机器人训练数据
在大规模数据集(如 Open X-Embodiment, OXE)上训练的通用机器人策略虽在多种任务中表现良好,但泛化能力受限于训练数据分布。本文揭示其根本原因在于‘捷径学习’——即过度依赖任务无关特征。通过理论与实证分析,发现两大关键因素:(1) 单个子数据集内多样性有限;(2) 不同子数据集间存在显著分布差异,造成数据碎片化。这些现象源于像 OXE 这类数据集由不同环境、不同机器人形态独立采集而成的结构特性。研究结果为改进数据收集策略提供关键洞见,能有效减少捷径学习,提升泛化能力。此外,在无法获取新大规模数据时,本文证明精心设计的数据增强策略可有效改善现有离线数据集中的捷径学习问题,从而提升策略 $π_0$ 在仿真与真实环境中的泛化表现。
原文摘要 · Abstract (English)
Generalist robot policies trained on large-scale datasets such as Open X-Embodiment (OXE) demonstrate strong performance across a wide range of tasks. However, they often struggle to generalize beyond the distribution of their training data. In this paper, we investigate the underlying cause of this limited generalization capability. We identify shortcut learning -- the reliance on task-irrelevant features -- as a key impediment to generalization. Through comprehensive theoretical and empirical analysis, we uncover two primary contributors to shortcut learning: (1) limited diversity within individual sub-datasets, and (2) significant distributional disparities across sub-datasets, leading to dataset fragmentation. These issues arise from the inherent structure of large-scale datasets like OXE, which are typically composed of multiple sub-datasets collected independently across varied environments and embodiments. Our findings provide critical insights into dataset collection strategies that can reduce shortcut learning and enhance the generalization ability of generalist robot policies. Moreover, in scenarios where acquiring new large-scale data is impractical, we demonstrate that carefully selected robotic data augmentation strategies can effectively reduce shortcut learning in existing offline datasets, thereby improving generalization capabilities of generalist robot policies, e.g., $π_0$, in both simulation and real-world environments. More information at https://lucky-light-sun.github.io/proj/shortcut-learning-in-grps/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。