arXiv:2606.24799cs.CVcs.AI2026-06

用视频生成3D场景,让文字变出完整环绕的立体画面。

OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis

论文配图:OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis
图 1 · 摘自论文原文
  • 用重建锚定视频,补全缺失视角,生成闭环3D场景
  • 单视频生成覆盖达359度,图像奖励提升近一倍
  • 无需微调,适合快速生成高质量3D内容的用户

通用文本到视频模型可作为丰富开放世界场景先验。尽管当前生成视频质量高,但难以直接得到可靠3D资产:相机运动难控、视图覆盖不全、帧间常有不一致。我们提出OrbitForge,一种基于冻结视频先验与每提示词高斯点云重建优化的适配器,将单个文本生成视频转换为规范闭合轨道的3D高斯点云场景。以3D重建为锚点,提升生成视频的3D一致性。首先通过可变形高斯点云(Deformable Gaussian Splatting)和鲁棒中值代理(MedianGS)获取初步3D重建,再按预定轨道渲染视图,检测缺失视角。OrbitForge仅使用文本到视频模型补全缺失视图,并将完整轨道重建为最终高斯点云场景。该设计无需任务特定视频或多视图微调,避免逐帧生成或分数蒸馏优化。我们进一步主张应采用覆盖感知评估:仅局部平滑会奖励那些从未尝试完整轨道的方法。在冻结300提示词的T3Bench衍生审计中,OrbitForge重建达到359.0度的中位跨度,使原本不支持的Q10 ImageReward从8.07提升至16.36,同时在覆盖-质量上与VideoMV保持竞争力。

原文摘要 · Abstract (English)

Generic text-to-video models can be used as rich open-world scene priors. Despite the high quality of today's generated videos, they do not directly yield reliable 3D assets: camera motion is difficult to control, view coverage is partial, and frames often contain inconsistencies across time. We introduce OrbitForge, an adapter built from frozen video priors and per-prompt Gaussian Splatting reconstruction optimization that converts a single text-generated video into a canonical closed-orbit 3D Gaussian Splatting scene. We use 3D reconstruction as an anchor to improve the 3D consistency of the generated video. We obtain a preliminary 3D reconstruction from a first generated video via Deformable Gaussian Splatting with a robust MedianGS proxy. We render views from a prescribed orbit to detect missing viewpoints. OrbitForge uses the text-to-video model to complete only the missing views, and reconstructs the completed orbit into a final Gaussian Splatting scene. This design requires no task-specific video or multiview fine-tuning, avoids per-prompt score-distillation optimization, and does not progressively generate views one step at a time. We further argue that this setting demands coverage-aware evaluation: local smoothness alone rewards methods that never attempt a full orbit. On a frozen 300-prompt T3Bench-derived audit, OrbitForge reconstruction attains a 359.0-degree measured median span, raises originally unsupported-bin Q10 ImageReward from 8.07 to 16.36 relative to MedianGS-only reconstruction, while remaining competitive with VideoMV on the coverage-quality.

文本生成3D高斯点云视频重建多视角生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。