arXiv:2410.05259cs.CV2024-10IJCV被引 15

用3D高斯泼溅实现可控、连贯的3D虚拟试穿

GS-VTON: Controllable 3D Virtual Try-on with Gaussian Splatting

  • 以图像提示替代文本,通过低秩适配融合个性化信息
  • 多视角一致编辑,解决2D结果在3D中的外观不一致问题
  • 自建3D-VTONBench基准,支持全面评估

基于扩散模型的2D虚拟试穿技术表现优异,但3D VTON发展滞后。主要原因在于:文本提示难以描述服装细节,且同一3D场景不同视角生成的2D结果缺乏一致性,导致外观不连贯和几何畸变。为此,我们提出图像提示的3D VTON方法GS-VTON,利用3D高斯泼溅(3DGS)作为3D表示,将预训练2D VTON模型知识迁移至3D并提升跨视角一致性。(1)提出个性化扩散模型,采用低秩适配(LoRA)微调融合个性化信息;为有效训练LoRA,设计参考驱动的多视角图像编辑方法,确保同步编辑与一致性。(2)提出人物感知的3DGS编辑框架,实现高效编辑的同时保持跨视角外观一致性和高质量3D几何结构。(3)构建新基准3D-VTONBench,支持全面的定性与定量评估。大量实验表明,所提方法在保真度和编辑能力上均优于现有方法。

原文摘要 · Abstract (English)

Diffusion-based 2D virtual try-on (VTON) techniques have recently demonstrated strong performance, while the development of 3D VTON has largely lagged behind. Despite recent advances in text-guided 3D scene editing, integrating 2D VTON into these pipelines to achieve vivid 3D VTON remains challenging. The reasons are twofold. First, text prompts cannot provide sufficient details in describing clothing. Second, 2D VTON results generated from different viewpoints of the same 3D scene lack coherence and spatial relationships, hence frequently leading to appearance inconsistencies and geometric distortions. To resolve these problems, we introduce an image-prompted 3D VTON method (dubbed GS-VTON) which, by leveraging 3D Gaussian Splatting (3DGS) as the 3D representation, enables the transfer of pre-trained knowledge from 2D VTON models to 3D while improving cross-view consistency. (1) Specifically, we propose a personalized diffusion model that utilizes low-rank adaptation (LoRA) fine-tuning to incorporate personalized information into pre-trained 2D VTON models. To achieve effective LoRA training, we introduce a reference-driven image editing approach that enables the simultaneous editing of multi-view images while ensuring consistency. (2) Furthermore, we propose a persona-aware 3DGS editing framework to facilitate effective editing while maintaining consistent cross-view appearance and high-quality 3D geometry. (3) Additionally, we have established a new 3D VTON benchmark, 3D-VTONBench, which facilitates comprehensive qualitative and quantitative 3D VTON evaluations. Through extensive experiments and comparative analyses with existing methods, the proposed \OM has demonstrated superior fidelity and advanced editing capabilities, affirming its effectiveness for 3D VTON.

虚拟试穿3D高斯可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。