让个性化视频生成具备真实3D结构,支持多视角一致的动态展示。
3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model
- 用单帧优化分离几何与运动,注入强3D先验
- 结合多视角联合优化,实现细节清晰、收敛快的生成效果
- 适合虚拟人、VR/AR、电商等需3D定制视频的场景
为沉浸式VR/AR、虚拟制作及下一代电商等应用,构建可动态呈现的定制化主体视频至关重要。然而现有方法多将主体视为2D实体,依赖单视图视觉特征或文本提示传递身份,缺乏真实3D物体所需的完整空间先验。这导致在生成新视角时,只能凭空生成未见区域的细节,无法保持真实3D一致性。由于多视角视频数据集稀缺,直接在有限视频序列上微调易引发时间过拟合。为此,我们提出3DreamBooth与3Dapter协同框架:3DreamBooth通过1帧优化范式解耦空间几何与时间运动,仅更新空间表示,无需全视频训练即可嵌入稳健3D先验;3Dapter作为视觉条件模块,采用非对称条件策略与主干分支进行多视图联合优化,动态从极简参考集查询视图特定几何提示,加速收敛并增强纹理精细度。
原文摘要 · Abstract (English)
Creating dynamic, view-consistent videos of customized subjects is highly sought after for a wide range of emerging applications, including immersive VR/AR, virtual production, and next-generation e-commerce. However, despite rapid progress in subject-driven video generation, existing methods predominantly treat subjects as 2D entities, focusing on transferring identity through single-view visual features or textual prompts. Because real-world subjects are inherently 3D, applying these 2D-centric approaches to 3D object customization reveals a fundamental limitation: they lack the comprehensive spatial priors necessary to reconstruct the 3D geometry. Consequently, when synthesizing novel views, they must rely on generating plausible but arbitrary details for unseen regions, rather than preserving the true 3D identity. Achieving genuine 3D-aware customization remains challenging due to the scarcity of multi-view video datasets. While one might attempt to fine-tune models on limited video sequences, this often leads to temporal overfitting. To resolve these issues, we introduce a novel framework for 3D-aware video customization, comprising 3DreamBooth and 3Dapter. 3DreamBooth decouples spatial geometry from temporal motion through a 1-frame optimization paradigm. By restricting updates to spatial representations, it effectively bakes a robust 3D prior into the model without the need for exhaustive video-based training. To enhance fine-grained textures and accelerate convergence, we incorporate 3Dapter, a visual conditioning module. Following single-view pre-training, 3Dapter undergoes multi-view joint optimization with the main generation branch via an asymmetrical conditioning strategy. This design allows the module to act as a dynamic selective router, querying view-specific geometric hints from a minimal reference set. Project page: https://ko-lani.github.io/3DreamBooth/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。