arXiv:2608.20759cs.CV2026-08

用扩散模型快速生成可动画3D人体,0.71秒完成

DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion

论文配图:DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion
图 1 · 摘自论文原文
  • 将单图重建转为UV空间扩散补全,保证3D一致性
  • 教师-学生框架生成几何对齐伪真值,提升结构精度
  • 支持零样本泛化,生成结果可直接用于动画

单图像3D人体重建常面临纹理过平滑和几何不一致问题。尽管扩散模型提升了生成质量,但其依赖多视角合成前驱,计算成本高且易产生视图不一致。我们提出DiGS-Avatar,将该任务重构为高效的基于扩散的UV空间潜在补全,通过设计确保3D一致性。为捕捉精确空间结构,引入教师-学生框架:多视角教师提供几何对齐的伪真值潜在表示,监督单视角扩散学生;将推断出的潜在表示视为稳健的结构骨架,注入高层语义特征以精准恢复细微纹理细节,同时不破坏空间完整性。最终将优化后的表示解码为3D高斯原语。大量实验表明,DiGS-Avatar在视觉保真度与零样本泛化能力上达到或接近当前最优水平,且仅需0.71秒即可重建出完全可动画的3D化身。代码已开源。

原文摘要 · Abstract (English)

Single-image 3D human reconstruction often suffers from over-smoothed textures and geometric inconsistencies. While diffusion models improve generative quality, their reliance on multi-view synthesis prior to 3D reconstruction is computationally expensive and prone to view inconsistency. We propose DiGS-Avatar, which reformulates this task as an efficient, diffusion-based UV-latent completion task, ensuring 3D consistency by design. To capture accurate spatial structure, we introduce a teacher-student framework where a multi-view teacher provides geometrically aligned pseudo-ground-truth latents to supervise a single-view diffusion student. Treating this inferred latent as a robust structural skeleton, our method injects high-level semantic features to accurately recover fine textural details without disrupting spatial integrity. The refined representation is then decoded into 3D Gaussian primitives. Extensive experiments demonstrate that DiGS-Avatar achieves state-of-the-art or highly competitive visual fidelity and zero-shot generalization, while reconstructing a fully animatable 3D avatar in just 0.71 seconds. Code is available at https://github.com/KLMAV-CUC/DiGS-Avatar.

3D重建扩散模型可动画单图生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。