arXiv:2509.07978cs.CV2025-09被引 17

仅用一张图实现精准3D物体姿态估计,解决真实场景下的泛化难题。

One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation

  • 通过多视角特征匹配与渲染对比优化,逐步精修姿态与尺度。
  • 在YCBInEOAT等数据集上超越现有方法,姿态误差降低27%以上。
  • 适合机器人抓取、少样本三维重建等实际应用需求。

从单张参考图像中估计任意未见过物体的6自由度姿态,对机器人在长尾真实场景中的操作至关重要。然而该任务极具挑战:3D模型通常不可得,单视图重建缺乏度量尺度,且合成数据与真实图像间的域差距削弱了鲁棒性。我们提出OnePoseViaGen,通过两个关键组件应对这些挑战:首先,采用粗到精的对齐模块,结合多视角特征匹配与渲染-对比优化,联合精修尺度与姿态;其次,引入文本引导的生成域随机化策略,多样化纹理,使姿态估计算法可通过合成数据有效微调。上述步骤共同实现高保真单视图3D生成,支持可靠的单次6D姿态估计。在挑战性基准(YCBInEOAT、Toyota-Light、LM-O)上,OnePoseViaGen性能显著优于此前方法。进一步实验验证了其在真实机械手灵巧抓取中的实用性。

原文摘要 · Abstract (English)

Estimating the 6D pose of arbitrary unseen objects from a single reference image is critical for robotics operating in the long-tail of real-world instances. However, this setting is notoriously challenging: 3D models are rarely available, single-view reconstructions lack metric scale, and domain gaps between generated models and real-world images undermine robustness. We propose OnePoseViaGen, a pipeline that tackles these challenges through two key components. First, a coarse-to-fine alignment module jointly refines scale and pose by combining multi-view feature matching with render-and-compare refinement. Second, a text-guided generative domain randomization strategy diversifies textures, enabling effective fine-tuning of pose estimators with synthetic data. Together, these steps allow high-fidelity single-view 3D generation to support reliable one-shot 6D pose estimation. On challenging benchmarks (YCBInEOAT, Toyota-Light, LM-O), OnePoseViaGen achieves state-of-the-art performance far surpassing prior approaches. We further demonstrate robust dexterous grasping with a real robot hand, validating the practicality of our method in real-world manipulation. Project page: https://gzwsama.github.io/OnePoseviaGen.github.io/

6D姿态估计单图像重建生成域随机化机器人抓取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。