arXiv:2505.08239cs.GRcs.CV2025-05被引 2

自适应相机轨迹让单视角3D重建更完整、更一致

ACT-R: Adaptive Camera Trajectories for Single View 3D Reconstruction

  • 根据物体形状动态规划相机路径,主动暴露被遮挡区域
  • 在GSO数据集上显著优于现有方法,重建精度大幅提升
  • 无需实时训练,仅靠预训练模型前向推理,效率高

我们提出将自适应视点规划引入多视角合成,以提升单视角3D重建中的遮挡揭示与3D一致性。不同于独立或同时生成无序视图,我们生成一个有序的相机序列,利用时间一致性增强3D连贯性。更重要的是,相机轨迹并非预先固定,而是通过计算自适应相机轨迹(ACT)形成轨道,旨在最大化待重建3D物体被遮挡区域的可见性。一旦找到最优轨道,便将其输入视频扩散模型生成轨道周围的全新视图,再交由任意多视角3D重建模型获得最终结果。该多视角合成流程高效,无需运行时训练或优化,仅需调用预训练模型进行前向推理。实验表明,本方法能有效预测揭示遮挡的相机轨迹,并生成一致的新视图,在未见的GSO数据集上显著超越当前最优方法。

原文摘要 · Abstract (English)

We introduce the simple idea of adaptive view planning to multi-view synthesis, aiming to improve both occlusion revelation and 3D consistency for single-view 3D reconstruction. Instead of producing an unordered set of views independently or simultaneously, we generate a sequence of views, leveraging temporal consistency to enhance 3D coherence. More importantly, our view sequence is not determined by a pre-determined and fixed camera setup. Instead, we compute an adaptive camera trajectory (ACT), forming an orbit, which seeks to maximize the visibility of occluded regions of the 3D object to be reconstructed. Once the best orbit is found, we feed it to a video diffusion model to generate novel views around the orbit, which can then be passed to any multi-view 3D reconstruction model to obtain the final result. Our multi-view synthesis pipeline is quite efficient since it involves no run-time training/optimization, only forward inferences by applying pre-trained models for occlusion analysis and multi-view synthesis. Our method predicts camera trajectories that reveal occlusions effectively and produce consistent novel views, significantly improving 3D reconstruction over SOTA alternatives on the unseen GSO dataset. Project Page: https://mingrui-zhao.github.io/ACT-R/

3D重建相机轨迹扩散模型视点规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。