用锚点策略解决机器人操作中数据多样性陷阱问题
Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation

- 先在关键场景重复演示稳定基础策略,再针对性扩展高风险边界
- 相同数据预算下,任务成功率提升显著,错误率更低
- 适合资源受限的实机部署场景,尤其对数据采集成本高的研究者
视觉-语言-动作(VLA)模型虽具强泛化能力,但部署于具体硬件时需克服具身差距。由于真实机器人示范成本高昂,适应过程常受限于严格的数据预算。本文发现:盲目追求多样性的单次示范策略可能因不可忽略的估计噪声而自陷困境,形成‘覆盖-密度权衡’现象。通过将策略误差分解为估计(密度)与外推(覆盖)两部分,我们揭示了固定预算下的最优独特条件分配。基于此,提出锚点中心适应(ACA)框架:第一阶段在核心锚点上通过重复示范稳定策略骨架;第二阶段通过教师强制误差挖掘与约束残差更新,选择性扩展至高风险边界。真实机器人实验验证了该权衡框架的有效性,证明在相同预算下,ACA 显著优于传统多样化采样策略,大幅提升任务可靠性与成功率。
原文摘要 · Abstract (English)
While Vision-Language-Action (VLA) models offer broad general capabilities, deploying them on specific hardware requires real-world adaptation to bridge the embodiment gap. Since robot demonstrations are costly, this adaptation must often occur under a strict data budget. In this work, we identify a critical diversity trap: the standard heuristic of "maximizing coverage" by collecting diverse, single-shot demonstrations can be self-defeating due to non-vanishing estimation noise. We formalize this phenomenon as a Coverage--Density Trade-off. By decomposing the policy error into estimation (density) and extrapolation (coverage) terms, we characterize an interior optimal allocation of unique conditions for a fixed budget. Guided by this analysis, we propose Anchor-Centric Adaptation (ACA), a two-stage framework that first stabilizes a policy skeleton through repeated demonstrations at core anchors, then selectively expands coverage to high-risk boundaries via teacher-forced error mining and constrained residual updates. Real-robot experiments validate our trade-off framework and demonstrate that ACA significantly improves task reliability and success rates over standard diverse sampling strategies under the same budget.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。