arXiv:2604.10789cs.CV2026-04被引 1

零样本将视频转为结构化3D场景,无需人工提示

ReplicateAnyScene: Zero-Shot Video-to-3D Composition via Textual-Visual-Spatial Alignment

  • 通过文本-视觉-空间对齐,从视频中自动提取通用先验
  • 在C3DR基准上生成高质量、物理合理的3D场景
  • 适合做空间智能和具身AI研究的开发者

人类具有从视频中快速感知并分割物体,并在脑中组装成结构化3D场景的天然能力。实现这种能力,称为组合式3D重建,对空间智能和具身AI的发展至关重要。然而,现有方法因跨模态信息整合不足,仍依赖人工物体提示、需要额外视觉输入,且受训练偏差限制,仅能处理过于简单的场景,难以实用。为此,我们提出ReplicateAnyScene框架,可全自动、零样本地将随意拍摄的视频转化为组合式3D场景。其五阶段级联流程从视觉基础模型中提取并结构化对齐文本、视觉与空间维度的通用先验,将其落地为结构化3D表示,确保场景语义一致性和物理合理性。为更全面评估该任务,我们进一步引入C3DR基准,从多维度衡量重建质量。大量实验表明,该方法在生成高质量组合式3D场景方面显著优于现有基线。

原文摘要 · Abstract (English)

Humans exhibit an innate capacity to rapidly perceive and segment objects from video observations, and even mentally assemble them into structured 3D scenes. Replicating such capability, termed compositional 3D reconstruction, is pivotal for the advancement of Spatial Intelligence and Embodied AI. However, existing methods struggle to achieve practical deployment due to the insufficient integration of cross-modal information, leaving them dependent on manual object prompting, reliant on auxiliary visual inputs, and restricted to overly simplistic scenes by training biases. To address these limitations, we propose ReplicateAnyScene, a framework capable of fully automated and zero-shot transformation of casually captured videos into compositional 3D scenes. Specifically, our pipeline incorporates a five-stage cascade to extract and structurally align generic priors from vision foundation models across textual, visual, and spatial dimensions, grounding them into structured 3D representations and ensuring semantic coherence and physical plausibility of the constructed scenes. To facilitate a more comprehensive evaluation of this task, we further introduce the C3DR benchmark to assess reconstruction quality from diverse aspects. Extensive experiments demonstrate the superiority of our method over existing baselines in generating high-quality compositional 3D scenes.

3D重建视频生成空间智能零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。