arXiv:2501.04144cs.CVcs.GR2025-01被引 2

用2D图像生成带精细结构的3D物体,支持部件自由组合。

Chirpy3D: Part-Aware Multi-View Diffusion for Creative Fine-Grained Object Generation

  • 基于2D图像和分割掩码构建分层部件潜在空间。
  • 实现部件级交换、插值与零样本组合生成。
  • 无需3D数据或人工标注,适合创意设计场景。

理解并生成物体的细粒度结构(如具有物种特异性喙、翅膀和尾羽的鸟类)是计算机视觉中的长期挑战。我们提出Chirpy3D,一种部件感知的多视角扩散框架,仅使用现成的2D部件分割掩码作为空间引导,从无姿态的2D图像中学习分层部件潜在空间——无需任何3D数据、相机位姿或手动部件标注。该潜在空间支持直观的部件级交换、插值与零样本组合。自监督特征一致性损失进一步促进跨视角结构对齐,使混合或未见过的部件组合也能生成连贯结果。核心贡献在于可控的部件感知潜在空间与多视角扩散模型。下游3D生成可通过任意可微渲染器(如NeRF)实现,但与主框架正交,使Chirpy3D成为在缺乏结构化3D数据时进行创意物体生成的灵活基础。代码已开源。

原文摘要 · Abstract (English)

Understanding and generating the fine-grained structure of objects -- such as birds with species-specific beaks, wings, and tails -- is a long-standing challenge in computer vision. We propose Chirpy3D, a part-aware multi-view diffusion framework that learns a hierarchical part latent space from unposed 2D images, using only off-the-shelf 2D part segmentation masks as spatial guidance -- without requiring any 3D data, camera poses, or manual part annotations. This latent space enables intuitive part-level swapping, interpolation, and zero-shot composition. A self-supervised feature consistency loss further encourages structural alignment across views, allowing coherent generation even with hybrid or unseen part combinations. Our core contribution is the controllable part-aware latent space and multi-view diffusion model. Downstream 3D generation is supported via any differentiable renderer such as NeRF but is orthogonal to the main framework, making Chirpy3D a flexible foundation for creative object generation in the absence of structured 3D data. Code is released at https://github.com/kamwoh/chirpy3d.

3D生成扩散模型部件感知创意设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。