arXiv:2508.13065cs.CV2025-08

用深度图引导扩散模型,实现身份不变的逼真人形重塑。

Odo: Depth-Guided Diffusion for Identity-Preserving Body Reshaping

  • 结合冻结UNet与控制网,用深度图指导形状变化。
  • 重建误差低至7.5mm,优于基线方法的13.6mm。
  • 适用于影视、游戏等需保持身份一致的场景。

人体形态编辑可实现对人物体形(如瘦、健壮、肥胖)的可控变换,同时保留姿态、身份、服饰和背景。与快速发展的姿态编辑不同,体形编辑仍处于探索阶段。现有方法多依赖3D可变形模型或图像扭曲,常因对齐误差和形变导致比例失真、纹理异常和背景不一致。核心瓶颈在于缺乏大规模公开数据集。本文构建首个包含18,573张图像、覆盖1523名受试者的大型数据集,专为可控人体形态编辑设计,涵盖脂肪、肌肉、瘦削等多种体形,且在身份、服饰和背景上保持一致。基于此,我们提出Odo,一种端到端的扩散模型方法,通过简单语义属性实现逼真直观的体形重塑。该方法利用冻结的UNet保留输入图像的细节与背景,由控制网根据目标SMPL深度图引导形变。大量实验表明,本方法显著优于先前方法,每顶点重建误差低至7.5mm,远低于基线的13.6mm,且生成结果真实贴合目标形态。

原文摘要 · Abstract (English)

Human shape editing enables controllable transformation of a person's body shape, such as thin, muscular, or overweight, while preserving pose, identity, clothing, and background. Unlike human pose editing, which has advanced rapidly, shape editing remains relatively under-explored. Current approaches typically rely on 3D morphable models or image warping, often introducing unrealistic body proportions, texture distortions, and background inconsistencies due to alignment errors and deformations. A key limitation is the lack of large-scale, publicly available datasets for training and evaluating body shape manipulation methods. In this work, we introduce the first large-scale dataset of 18,573 images across 1523 subjects, specifically designed for controlled human shape editing. It features diverse variations in body shape, including fat, muscular and thin, captured under consistent identity, clothing, and background conditions. Using this dataset, we propose Odo, an end-to-end diffusion-based method that enables realistic and intuitive body reshaping guided by simple semantic attributes. Our approach combines a frozen UNet that preserves fine-grained appearance and background details from the input image with a ControlNet that guides shape transformation using target SMPL depth maps. Extensive experiments demonstrate that our method outperforms prior approaches, achieving per-vertex reconstruction errors as low as 7.5mm, significantly lower than the 13.6mm observed in baseline methods, while producing realistic results that accurately match the desired target shapes.

体形编辑扩散模型深度引导身份保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。