arXiv:2601.16982cs.CVcs.LG2026-01被引 7

AnyView可零样本生成任意视角的动态场景视频。

AnyView: Synthesizing Any Novel View in Dynamic Scenes

  • 基于扩散模型,融合多源数据训练通用时空隐式表征。
  • 在极端动态场景下仍能保持视频时空一致性,优于现有方法。
  • 适合需要任意视角生成的虚拟拍摄、元宇宙等应用。

现代生成式视频模型在高质量输出方面表现优异,但在高度动态的真实环境中难以维持多视角与时空一致性。本文提出AnyView,一种基于扩散模型的动态视图合成框架,无需强先验或几何假设。通过利用单目(2D)、多视角静态(3D)和多视角动态(4D)等多种数据源进行训练,构建具备通用性的时空隐式表示,可从任意相机位置和轨迹实现零样本新视频生成。我们在标准基准上评估AnyView,结果达到当前最优水平,并提出了新的挑战性基准AnyViewBench,专用于极端动态场景下的视图合成。在更剧烈的设定下,多数基线性能显著下降,因依赖视角间大量重叠;而AnyView仍能生成真实、合理且时空一致的视频,无论输入视角如何。相关结果、数据、代码与模型详见:https://tri-ml.github.io/AnyView/

原文摘要 · Abstract (English)

Modern generative video models excel at producing convincing, high-quality outputs, but struggle to maintain multi-view and spatiotemporal consistency in highly dynamic real-world environments. In this work, we introduce \textbf{AnyView}, a diffusion-based video generation framework for \emph{dynamic view synthesis} with minimal inductive biases or geometric assumptions. We leverage multiple data sources with various levels of supervision, including monocular (2D), multi-view static (3D) and multi-view dynamic (4D) datasets, to train a generalist spatiotemporal implicit representation capable of producing zero-shot novel videos from arbitrary camera locations and trajectories. We evaluate AnyView on standard benchmarks, showing competitive results with the current state of the art, and propose \textbf{AnyViewBench}, a challenging new benchmark tailored towards \emph{extreme} dynamic view synthesis in diverse real-world scenarios. In this more dramatic setting, we find that most baselines drastically degrade in performance, as they require significant overlap between viewpoints, while AnyView maintains the ability to produce realistic, plausible, and spatiotemporally consistent videos when prompted from \emph{any} viewpoint. Results, data, code, and models can be viewed at: https://tri-ml.github.io/AnyView/

视频生成扩散模型动态视图合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。