arXiv:2508.20470cs.CV2025-08被引 2

用视频中的常识先验提升3D生成的合理性与一致性

Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation

  • 从多视角视频中提取空间一致性先验,指导3D生成
  • 在400万条视频数据上训练,实现图文双输入3D生成
  • 适合需要真实感3D内容的创作者和工业级应用

规模化模型在文本、图像和视频生成领域已证明成功,但3D领域因数据稀缺而受限。相比之下,互联网上存在大量包含常识先验的多视角视频,可作为替代监督信号缓解泛化瓶颈。视频既提供物体或场景的多视图信息以建立空间一致性,又蕴含丰富语义,使生成内容更贴合文本提示且语义合理。本文提出Droplet3D-4M,首个具有多视图标注的大规模视频数据集,并训练了支持图像与密集文本输入的生成模型Droplet3D。实验验证其生成内容在空间一致性和语义合理性上表现优异。相较于主流3D方法,本方案具备扩展至场景级应用的潜力。结果表明,视频中的常识先验显著促进3D创作。所有资源(数据、代码、框架、权重)已开源:https://dropletx.github.io/。

原文摘要 · Abstract (English)

Scaling laws have validated the success and promise of large-data-trained models in creative generation across text, image, and video domains. However, this paradigm faces data scarcity in the 3D domain, as there is far less of it available on the internet compared to the aforementioned modalities. Fortunately, there exist adequate videos that inherently contain commonsense priors, offering an alternative supervisory signal to mitigate the generalization bottleneck caused by limited native 3D data. On the one hand, videos capturing multiple views of an object or scene provide a spatial consistency prior for 3D generation. On the other hand, the rich semantic information contained within the videos enables the generated content to be more faithful to the text prompts and semantically plausible. This paper explores how to apply the video modality in 3D asset generation, spanning datasets to models. We introduce Droplet3D-4M, the first large-scale video dataset with multi-view level annotations, and train Droplet3D, a generative model supporting both image and dense text input. Extensive experiments validate the effectiveness of our approach, demonstrating its ability to produce spatially consistent and semantically plausible content. Moreover, in contrast to the prevailing 3D solutions, our approach exhibits the potential for extension to scene-level applications. This indicates that the commonsense priors from the videos significantly facilitate 3D creation. We have open-sourced all resources including the dataset, code, technical framework, and model weights: https://dropletx.github.io/.

3D生成视频先验多视图常识推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。