arXiv:2512.10940cs.CVcs.AI2025-12被引 5

一个统一模型搞定3D/4D视角合成,支持多输入、可控制相机运动。

OmniView: An All-Seeing Diffusion Model for 3D and 4D View Synthesis

  • 分离表示空间、时间与视角条件,灵活组合输入
  • 在多视角新视图合成等任务上提升33%~60%图像质量
  • 适合需要统一生成框架的视觉生成研究者

现有方法将相机控制注入扩散模型时,仅针对特定4D一致性任务(如新视图合成、带相机控制的文本到视频等),导致训练数据碎片化。本文提出OmniView,一个覆盖广泛4D一致性任务的统一框架。该方法分别建模空间、时间和视角条件,支持灵活组合输入:可从静态、动态或多视角输入生成新视图,可反向或正向外推轨迹,也能根据文本或图像提示生成带完整相机控制的视频。OmniView在多个基准上表现媲美专用模型,在多视角新视图合成LLFF数据集上图像质量提升最高达33%,动态3D视频任务中提升60%,静态相机控制在RE-10K上提升20%,文本控制视频生成中相机轨迹误差降低4倍。该模型展示了通用4D视频生成的可行性。项目页面见https://snap-research.github.io/OmniView/

原文摘要 · Abstract (English)

Prior approaches injecting camera control into diffusion models have focused on specific subsets of 4D consistency tasks: novel view synthesis, text-to-video with camera control, image-to-video, amongst others. Therefore, these fragmented approaches are trained on disjoint slices of available 3D/4D data. We introduce OmniView, a unified framework that generalizes across a wide range of 4D consistency tasks. Our method separately represents space, time, and view conditions, enabling flexible combinations of these inputs. For example, OmniView can synthesize novel views from static, dynamic, and multiview inputs, extrapolate trajectories forward and backward in time, and create videos from text or image prompts with full camera control. OmniView is competitive with task-specific models across diverse benchmarks and metrics, improving image quality scores among camera-conditioned diffusion models by up to 33\% in multiview NVS LLFF dataset, 60\% in dynamic NVS Neural 3D Video benchmark, 20\% in static camera control on RE-10K, and reducing camera trajectory errors by 4x in text-conditioned video generation. With strong generalizability in one model, OmniView demonstrates the feasibility of a generalist 4D video model. Project page is available at https://snap-research.github.io/OmniView/

扩散模型3D生成视频生成相机控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。