arXiv:2508.06032cs.CV2025-08AAAI

用3D纹理生成模型提升人体衣物与部位的细粒度分割效果。

Learning 3D Texture-Aware Representations for Parsing Diverse Human Clothing and Body Parts

  • 利用图像到3D纹理扩散模型提取人体部位特征,实现精准对齐。
  • 支持任意数量人和未见服装类别的零样本分割,准确率显著提升。
  • 适合需要细粒度人体解析的应用,如虚拟试衣、动作分析。

现有方法在人体解析中常使用固定类别掩码,标签宽泛,难以区分细粒度衣物类型。近期开放词汇分割方法虽借助预训练文本到图像(T2I)扩散模型实现强零样本迁移,但通常将整个身体视为单一人物类别,无法区分多样衣物或详细身体部位。为此,我们提出Spectrum,一个统一网络,实现部件级像素分割(身体部位与衣物)和实例级分组。尽管基于扩散的开放词汇模型任务泛化能力强,其内部表示并不专用于精细人体解析。我们观察到,与具有宽泛表示的扩散模型不同,图像驱动的3D纹理生成器能保持与输入图像的忠实对应,从而生成更优的解析表示。Spectrum创新性地复用一个通过在3D人体纹理图上微调的T2I模型得到的图像到纹理(I2Tx)扩散模型,以增强与身体部位和衣物的对齐能力。从输入图像出发,通过I2Tx扩散模型提取人体内部特征,并在提示引导下生成语义有效且对齐多样衣物类别的掩码。训练完成后,Spectrum可为场景中任意数量的人生成每个可见身体部位和衣物类别的语义分割图,忽略孤立衣物或无关物体。我们在多个数据集上进行了广泛实验,分别评估身体部位、衣物部位、未见衣物类别及完整身体掩码,结果表明Spectrum在提示引导分割任务中持续优于基线方法。

原文摘要 · Abstract (English)

Existing methods for human parsing into body parts and clothing often use fixed mask categories with broad labels that obscure fine-grained clothing types. Recent open-vocabulary segmentation approaches leverage pretrained text-to-image (T2I) diffusion model features for strong zero-shot transfer, but typically group entire humans into a single person category, failing to distinguish diverse clothing or detailed body parts. To address this, we propose Spectrum, a unified network for part-level pixel parsing (body parts and clothing) and instance-level grouping. While diffusion-based open-vocabulary models generalize well across tasks, their internal representations are not specialized for detailed human parsing. We observe that, unlike diffusion models with broad representations, image-driven 3D texture generators maintain faithful correspondence to input images, enabling stronger representations for parsing diverse clothing and body parts. Spectrum introduces a novel repurposing of an Image-to-Texture (I2Tx) diffusion model (obtained by fine-tuning a T2I model on 3D human texture maps) for improved alignment with body parts and clothing. From an input image, we extract human-part internal features via the I2Tx diffusion model and generate semantically valid masks aligned to diverse clothing categories through prompt-guided grounding. Once trained, Spectrum produces semantic segmentation maps for every visible body part and clothing category, ignoring standalone garments or irrelevant objects, for any number of humans in the scene. We conduct extensive cross-dataset experiments, separately assessing body parts, clothing parts, unseen clothing categories, and full-body masks, and demonstrate that Spectrum consistently outperforms baseline methods in prompt-based segmentation.

人体解析细粒度分割扩散模型3D纹理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。