arXiv:2607.03891cs.CV2026-07中稿 · ECCV

用少量3D模板和2D图像实现高精度可动物体三维重建

BAT3R: Bootstrapping Articulated 3D Reconstruction from 2D Image Collections

论文配图:BAT3R: Bootstrapping Articulated 3D Reconstruction from 2D Image Collections
图 1 · 摘自论文原文
  • 基于弱监督的迭代优化框架,从单个基准网格开始逐步学习关节姿态
  • 仅需每类一个绑定好的原始网格,训练性能接近需人工标注数据的方法
  • 适合缺乏真实3D标注数据但有大量2D图像的场景

从单张图像进行可动物体的3D重建极具挑战,因配对的图像与3D标注数据难以获取。现有基于点图的方法虽表现优异,但依赖于手工创建的可动3D资产及精心设计的姿态分布所生成的合成数据。尽管相机视角易于采样,生成真实感的物体运动仍成本高昂且费力。我们提出一种训练框架,通过仅使用每类一个带绑定的基准网格和未标注的2D图像集合,显著降低对3D监督的需求。从在基准姿态渲染图像上训练的弱3D形状预测器出发,通过拟合预测点图迭代估计物体关节姿态与相机位姿,再利用恢复出的姿态和视角生成更新的合成训练数据,逐步提升预测器性能。尽管采用更弱的3D监督,我们的模型在性能上仍可媲美需要人工精细标注的DualPM方法。

原文摘要 · Abstract (English)

3D reconstruction of articulated objects from a single image is challenging because large training datasets with paired image and 3D supervision are difficult to obtain. Recent point map-based methods achieve strong performance but rely on synthetic datasets rendered from manually created articulated 3D assets with carefully curated pose distributions. While camera viewpoints can be easily sampled, generating realistic object articulations remains costly and labor-intensive. We propose a training framework that reduces this requirement by leveraging unannotated 2D images collections with only a single rigged canonical mesh per category. Starting from a weak 3D shape predictor trained on canonical-pose renders, we iteratively estimate object articulation and camera pose by fitting the mesh to predicted point maps. The recovered articulations and viewpoints are then used to render updated synthetic training data, progressively improving the predictor. Despite using substantially weaker 3D supervision, our models achieve performance comparable with DualPM, which requires manually curated articulated training datasets.

3D重建可动物体弱监督点图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。