用机器人成功经验生成新物体的逼真多视角数据,提升泛化能力。
Pose6DAug: Physically Plausible Multi-view Object Swapping for Robot Data Augmentation

- 基于成功轨迹在3D中替换目标物体,保持物理合理性。
- 在新物体上提升相对16.5%的成功率,不损失原有性能。
- 适合需要快速扩展泛化能力的机器人视觉-语言-动作系统。
视觉-语言-动作(VLA)策略在通用操作中展现巨大潜力,但对训练分布外的新物体常失效。传统方法需为每个失败案例重新采集多视角遥操作数据,成本高昂。我们提出Pose6DAug,一种故障驱动的数据增强框架,利用策略自身成功回合生成针对失败模式的目标示范,无需新数据采集。核心思路是:每个成功回合已包含物理合理的动作轨迹与校准的多视角观测。通过仅替换被操作物体并保留该轨迹,可生成新且物理成立的示范。然而,直接2D视频编辑会破坏多视角一致性与物理合理性,尤其在重遮挡和第一人称视角下。我们的方法在3D中进行,以时序一致的6D姿态轨迹驱动显式网格锚定目标物体,确保所有相机视角的几何一致性渲染。在增强数据上微调VLA,在新物体上的成功率相比当前最优基线提升16.5%,同时保持分布内性能。结果表明,多视角且物理一致的增强是实现可扩展VLA泛化的有效路径。
原文摘要 · Abstract (English)
Vision-language-action (VLA) policies have shown strong potential for general-purpose manipulation, yet they often fail on novel, out-of-distribution objects whose appearance or geometry deviates from the training distribution. The standard remedy is to collect multi-view teleoperation data for every failure case, but this scales poorly in both cost and time. We introduce Pose6DAug, a failure-driven data augmentation framework that turns a policy's own successful episodes into targeted demonstrations for its failure modes, without any new data collection. Our key insight is that each successful episode already encodes a physically valid action trajectory together with calibrated multi-view observations. By swapping only the manipulated object while preserving this trajectory, we obtain new and physically grounded demonstrations. However, naive 2D video editing breaks multi-view consistency and physical plausibility, particularly under heavy occlusion and egocentric viewpoints. Our method instead operates directly in 3D, anchoring the target object with an explicit mesh driven by a temporally coherent 6D pose trajectory, ensuring geometrically consistent renderings across all camera views. Fine-tuning a VLA on data augmented by our method improves success rates by 16.5% relative to the state-of-the-art baseline on novel objects, while preserving in-distribution performance. These results show that multi-view and physically consistent augmentation is a practical path to scalable VLA generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。