arXiv:2412.11525cs.CV2024-12被引 3

用视频超分模型提升3D重建一致性,简单对齐即达顶尖效果

Sequence Matters: Harnessing Video Models in 3D Super-Resolution

  • 基于视频超分模型,利用时序信息增强多视角一致性
  • 无需微调,在无精确对齐序列上仍实现顶级3D超分辨率性能
  • 方法极简但有效,适合追求高效高质3D重建的研究者

3D超分辨率旨在从低分辨率多视角图像重建高保真3D模型。早期方法主要依赖单图超分(SISR)模型将低分辨率图像上采样为高分辨率图像,但各图像独立处理导致视角间一致性差。尽管已探索多种后处理技术缓解该问题,仍未能彻底解决。本文通过引入视频超分(VSR)模型进行3D超分辨率研究,利用其时序建模能力保证空间一致性,并参考邻近帧信息以获得更准确、更精细的重建结果。实验发现,即使在缺乏精确空间对齐的序列上,VSR模型依然表现优异。据此,我们提出一种无需微调、不依赖生成平滑轨迹的简单对齐方法。在NeRF-synthetic和MipNeRF-360等标准数据集上的实验表明,该方法达到当前3D超分辨率任务的最先进水平。

原文摘要 · Abstract (English)

3D super-resolution aims to reconstruct high-fidelity 3D models from low-resolution (LR) multi-view images. Early studies primarily focused on single-image super-resolution (SISR) models to upsample LR images into high-resolution images. However, these methods often lack view consistency because they operate independently on each image. Although various post-processing techniques have been extensively explored to mitigate these inconsistencies, they have yet to fully resolve the issues. In this paper, we perform a comprehensive study of 3D super-resolution by leveraging video super-resolution (VSR) models. By utilizing VSR models, we ensure a higher degree of spatial consistency and can reference surrounding spatial information, leading to more accurate and detailed reconstructions. Our findings reveal that VSR models can perform remarkably well even on sequences that lack precise spatial alignment. Given this observation, we propose a simple yet practical approach to align LR images without involving fine-tuning or generating 'smooth' trajectory from the trained 3D models over LR images. The experimental results show that the surprisingly simple algorithms can achieve the state-of-the-art results of 3D super-resolution tasks on standard benchmark datasets, such as the NeRF-synthetic and MipNeRF-360 datasets. Project page: https://ko-lani.github.io/Sequence-Matters

3D重建视频超分多视角一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。