arXiv:2604.19747cs.CV2026-04被引 2

用视频扩散模型实现任意视角3D重建,支持多视角灵活输入。

AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model

论文配图:AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model
图 1 · 摘自论文原文
  • 构建全局场景记忆,支持多帧长期条件输入。
  • 在大视角变化下保持帧级对应,重建更稳定。
  • 适合复杂场景和长轨迹重建,效率高。

稀疏视角3D重建对从随意拍摄中建模场景至关重要,但传统非生成方法仍具挑战。现有基于扩散模型的方法虽能合成新视角,但通常仅依赖一至两帧输入,限制几何一致性并难以扩展至大规模或多样化场景。本文提出AnyRecon,一种可扩展的框架,支持任意、无序稀疏输入的3D重建,同时保留显式几何控制并支持灵活的条件输入数量。为实现长距离条件感知,方法通过预置捕获帧缓存构建持久的全局场景记忆,并去除时间压缩以维持大视角变化下的帧级对应关系。此外,我们发现生成与重建之间的相互作用对大规模场景至关重要,因此引入一种几何感知的条件策略,通过显式3D几何记忆和基于几何的捕获帧检索,耦合生成与重建过程。为保证效率,结合4步扩散蒸馏与上下文窗口稀疏注意力,将复杂度从二次降低。大量实验表明,该方法在不规则输入、大视角间隔及长轨迹下均能实现鲁棒且可扩展的重建。

原文摘要 · Abstract (English)

Sparse-view 3D reconstruction is essential for modeling scenes from casual captures, but remain challenging for non-generative reconstruction. Existing diffusion-based approaches mitigates this issues by synthesizing novel views, but they often condition on only one or two capture frames, which restricts geometric consistency and limits scalability to large or diverse scenes. We propose AnyRecon, a scalable framework for reconstruction from arbitrary and unordered sparse inputs that preserves explicit geometric control while supporting flexible conditioning cardinality. To support long-range conditioning, our method constructs a persistent global scene memory via a prepended capture view cache, and removes temporal compression to maintain frame-level correspondence under large viewpoint changes. Beyond better generative model, we also find that the interplay between generation and reconstruction is crucial for large-scale 3D scenes. Thus, we introduce a geometry-aware conditioning strategy that couples generation and reconstruction through an explicit 3D geometric memory and geometry-driven capture-view retrieval. To ensure efficiency, we combine 4-step diffusion distillation with context-window sparse attention to reduce quadratic complexity. Extensive experiments demonstrate robust and scalable reconstruction across irregular inputs, large viewpoint gaps, and long trajectories.

3D重建视频生成扩散模型几何控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。