通过诊断语言指令偏见,提升机器人对复杂指令的泛化能力
Scale Up Strategically: Learning Compositional Generalization via Bias-Aware Evaluation and Data Collection for Robotic Manipulation

- 定位指令各成分(如颜色、动作)的依赖程度,识别模型偷懒行为
- 发现颜色最被依赖,动词和尺寸最被忽视,偏差可量化为两个指标
- 按偏差调整数据收集策略,用一半演示数据实现更好性能
组合泛化对机器人理解多样化指令至关重要。然而预训练策略常走捷径,依赖显著线索而非真正理解语言。本文提出诊断框架,定位模型在各个指令因素(如颜色、动词、物体、大小、空间属性)上的依赖偏差。通过两个指标:因子主导率(FDR)衡量因子间成对偏差,因子主导层级(FDH)生成全局排序。在六个基础策略上评估显示一致结果:颜色 ≥ 物体 ≥ 空间 ≥ 动词 ≥ 大小,其中颜色主导,动词与大小最被忽视。该诊断可指导实践:采用偏差感知的数据收集策略,在固定预算下优先覆盖被忽视因素,显著提升模拟环境与真实机器人上的性能,仅需一半演示数据即优于基线,实现更高效、更强泛化的策略学习。
原文摘要 · Abstract (English)
Compositional generalization is essential for robot to follow diverse instructions. However, pretrained policies are known to take shortcuts, deferring to salient cues rather than grounding language. We introduce a diagnostic framework that localizes this failure to individual \textit{instruction factors}, \textit{e.g.,} reusable semantic components such as color, verb, object, size, and spatial attribute. Our framework formalizes instruction factor bias, the tendency of fine-tuned policies to over-rely on dominant factors as shortcuts, and quantifies it through two metrics: Factor Dominance Rate (FDR), capturing pairwise bias between factors, and Factor Dominance Hierarchy (FDH), aggregating these into a global ranking. Evaluation on six foundation policies reveals broadly consistent ordering, \textit{i.e.}, color $\geq$ object $\geq$ spatial $\geq$ verb $\geq$ size, with color dominant, and verb and size most under-grounded. We further show the diagnosis is actionable: a bias-aware data collection strategy that reallocates a fixed budget toward under-grounded factors outperforms baselines in simulation and on a real robot using half the demonstrations, thereby enabling more sample-efficient and generalizable policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。