arXiv:2509.24797cs.ROcs.AI2025-09被引 3

通过感知数据保真度,提升机器人在未知场景下的泛化能力。

Fidelity-Aware Data Composition for Robust Robot Generalization

  • 将数据组合建模为优化问题,以保真度为导向
  • 在真实与合成数据混合时,提升54%以上未知场景成功率
  • 适合关注机器人鲁棒性与泛化能力的研究者

在大规模、视觉同质的数据集上训练的通用机器人策略容易产生捷径学习,影响其在分布外(OOD)场景下的泛化能力。尽管生成式数据增强是引入多样性的常见方法,但数据组合存在隐性挑战:盲目混合真实与合成数据会因牺牲信息保真度而污染学习信号。本文提出,鲁棒泛化依赖于有原则的保真度感知数据组合。我们引入相干信息保真度调优(CIFT),将数据组合视为优化问题。CIFT基于数据集特征空间几何设计了信息保真度的实用代理指标,识别出训练稳定性下降的相变点——退相干点。框架包含多视角视频增强(MVAug)生成因果解耦的数据谱用于调优。应用于π₀和扩散策略(Diffusion Policy)等架构,在分布外场景的成功率提升超过54%。结果表明,仅靠数据合成不足以实现鲁棒性,保真度感知组合是构建通用机器人的关键。

原文摘要 · Abstract (English)

Generalist robot policies trained on large-scale, visually homogeneous datasets can be susceptible to shortcut learning, which impairs their out-of-distribution (OOD) generalization. While generative data augmentation is a common approach to introduce diversity, it presents a subtle challenge: data composition. Naively mixing real and synthetic data can corrupt the learning signal, as this process often prioritizes visual diversity at the expense of information fidelity. This paper suggests that robust generalization depends on principled, fidelity-aware data composition. We introduce Coherent Information Fidelity Tuning (CIFT), a framework that treats data composition as an optimization problem. CIFT uses a practical proxy for Information Fidelity based on the feature-space geometry of a dataset. This enables the identification of a phase transition, termed the Decoherence Point, where training stability degrades. The framework includes a generative engine, Multi-View Video Augmentation (MVAug), to synthesize a causally disentangled data spectrum for this tuning process. Applying CIFT to policy architectures such as $π_0$ and Diffusion Policy improves OOD success rates by over 54\%. These results indicate that fidelity-aware composition, beyond data synthesis alone, is an important component for developing robust, general-purpose robots.

机器人泛化数据组合保真度扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。