arXiv:2504.11427cs.CV2025-04ICCV被引 24

利用视频扩散模型先验,提升视频法向量估计的时序一致性。

NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion Priors

  • 基于扩散模型的时序先验,结合语义特征正则化对齐场景本质
  • 两阶段训练在潜空间与像素空间协同学习,保持空间精度与长时上下文
  • 在多类视频上生成细节丰富、时序一致的法向量序列

表面法向量估计是众多计算机视觉任务的基础。尽管静态图像场景已有大量研究,但视频中法向量估计的时序一致性仍是难题。我们提出NormalCrafter,不简单叠加时序模块,而是利用视频扩散模型固有的时序先验。为确保序列中高保真法向量估计,提出语义特征正则化(SFR),使扩散特征与语义线索对齐,促使模型聚焦于场景内在语义。此外,设计两阶段训练策略,融合潜空间与像素空间学习,在保持空间精度的同时维持长时上下文。大量实验表明,该方法能从多样视频中生成细节丰富、时序一致的法向量序列,性能显著优于现有方法。

原文摘要 · Abstract (English)

Surface normal estimation serves as a cornerstone for a spectrum of computer vision applications. While numerous efforts have been devoted to static image scenarios, ensuring temporal coherence in video-based normal estimation remains a formidable challenge. Instead of merely augmenting existing methods with temporal components, we present NormalCrafter to leverage the inherent temporal priors of video diffusion models. To secure high-fidelity normal estimation across sequences, we propose Semantic Feature Regularization (SFR), which aligns diffusion features with semantic cues, encouraging the model to concentrate on the intrinsic semantics of the scene. Moreover, we introduce a two-stage training protocol that leverages both latent and pixel space learning to preserve spatial accuracy while maintaining long temporal context. Extensive evaluations demonstrate the efficacy of our method, showcasing a superior performance in generating temporally consistent normal sequences with intricate details from diverse videos.

法向量估计视频生成扩散模型时序一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。