用文字提示增强草图语义,实现更精准的3D形状生成
PASTA: Part-Aware Sketch-to-3D Shape Generation with Text-Aligned Prior
- 结合草图与文本,用视觉语言模型提升草图语义
- 引入两种图卷积网络,分别处理细节与部件结构
- 支持细粒度部件编辑,适合需要精确控制的生成场景
条件3D形状生成的核心挑战在于最小化用户输入的信息损失并最大化其意图表达。现有方法主要依赖孤立的草图或文本描述,难以灵活控制生成结果。本文提出PASTA,一种将用户草图与文本描述无缝融合的灵活生成方法。核心思路是利用视觉语言模型的文本嵌入,为草图注入语义先验,明确物体部件组成,弥补模糊草图缺失的视觉线索。同时,提出ISG-Net,包含两种图卷积网络:IndivGCN处理细粒度细节,PartGCN将细节聚合为部件并优化整体结构。大量实验表明,PASTA在部件级编辑能力上优于现有方法,并在草图到3D形状生成任务中达到当前最优性能。
原文摘要 · Abstract (English)
A fundamental challenge in conditional 3D shape generation is to minimize the information loss and maximize the intention of user input. Existing approaches have predominantly focused on two types of isolated conditional signals, i.e., user sketches and text descriptions, each of which does not offer flexible control of the generated shape. In this paper, we introduce PASTA, the flexible approach that seamlessly integrates a user sketch and a text description for 3D shape generation. The key idea is to use text embeddings from a vision-language model to enrich the semantic representation of sketches. Specifically, these text-derived priors specify the part components of the object, compensating for missing visual cues from ambiguous sketches. In addition, we introduce ISG-Net which employs two types of graph convolutional networks: IndivGCN, which processes fine-grained details, and PartGCN, which aggregates these details into parts and refines the structure of objects. Extensive experiments demonstrate that PASTA outperforms existing methods in part-level editing and achieves state-of-the-art results in sketch-to-3D shape generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。