arXiv:2412.15491cs.CV2024-12被引 6

无需生成数据,一键适配3D生成模型到新场景。

GCA-3D: Towards Generalized and Consistent Domain Adaptation of 3D Generators

  • 用多模态深度感知损失直接优化3D生成器,跳过繁琐的数据生成流程。
  • 支持单图引导适配,保持姿态与身份一致性,效果优于现有方法。
  • 适合需要快速迁移3D生成模型的研究者和开发者。

近期,3D生成领域自适应方法旨在不依赖大规模数据集和相机位姿分布的情况下,将预训练生成器适配至新域。传统方法通常利用大规模文本到图像扩散模型合成目标域图像,再微调3D模型,但其数据生成流程复杂,不可避免引入源域与合成数据间的姿态偏差。此外,它们无法支持更难的单图引导自适应,因单图参考会引入更严重的姿态偏差和身份偏差。为此,我们提出GCA-3D,一种无需复杂数据生成流程的通用且一致的3D领域自适应方法。不同于以往流水线方法,GCA-3D引入多模态深度感知得分蒸馏采样损失,在非对抗框架下高效适配3D生成模型。该损失使GCA-3D同时支持文本提示与单图提示适配。此外,利用体渲染模块生成的每实例深度图,缓解过拟合问题并保持结果多样性。为增强姿态与身份一致性,进一步提出分层空间一致性损失,对齐源域与目标域生成图像的空间结构。实验表明,GCA-3D在效率、泛化性、姿态准确率和身份一致性方面均优于现有方法。

原文摘要 · Abstract (English)

Recently, 3D generative domain adaptation has emerged to adapt the pre-trained generator to other domains without collecting massive datasets and camera pose distributions. Typically, they leverage large-scale pre-trained text-to-image diffusion models to synthesize images for the target domain and then fine-tune the 3D model. However, they suffer from the tedious pipeline of data generation, which inevitably introduces pose bias between the source domain and synthetic dataset. Furthermore, they are not generalized to support one-shot image-guided domain adaptation, which is more challenging due to the more severe pose bias and additional identity bias introduced by the single image reference. To address these issues, we propose GCA-3D, a generalized and consistent 3D domain adaptation method without the intricate pipeline of data generation. Different from previous pipeline methods, we introduce multi-modal depth-aware score distillation sampling loss to efficiently adapt 3D generative models in a non-adversarial manner. This multi-modal loss enables GCA-3D in both text prompt and one-shot image prompt adaptation. Besides, it leverages per-instance depth maps from the volume rendering module to mitigate the overfitting problem and retain the diversity of results. To enhance the pose and identity consistency, we further propose a hierarchical spatial consistency loss to align the spatial structure between the generated images in the source and target domain. Experiments demonstrate that GCA-3D outperforms previous methods in terms of efficiency, generalization, pose accuracy, and identity consistency.

3D生成领域自适应图像生成深度感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。