arXiv:2412.02912cs.CVcs.AI2024-12CVPR被引 4

用3D形状提示引导文生图,生成更符合描述的立体图像。

ShapeWords: Guiding Text-to-Image Synthesis with 3D Shape-Aware Prompts

  • 将3D形状信息嵌入特殊标记,与文本一起指导图像生成。
  • 生成图像在几何结构和文本描述上都更准确一致。
  • 适合需要精确形状控制的创意设计、工业建模场景。

我们提出ShapeWords,一种基于3D形状引导和文本提示的图像生成方法。ShapeWords将目标3D形状信息融入与输入文本共同嵌入的专用标记中,有效结合3D形状感知与文本上下文,指导图像合成过程。与传统依赖固定视角深度图、常忽略完整3D结构或文本上下文的形状引导方法不同,ShapeWords能生成多样且一致的图像,准确反映目标形状的几何特征与文本描述。实验表明,该方法生成的图像在文本契合度、美学合理性以及3D形状感知方面均有提升。

原文摘要 · Abstract (English)

We introduce ShapeWords, an approach for synthesizing images based on 3D shape guidance and text prompts. ShapeWords incorporates target 3D shape information within specialized tokens embedded together with the input text, effectively blending 3D shape awareness with textual context to guide the image synthesis process. Unlike conventional shape guidance methods that rely on depth maps restricted to fixed viewpoints and often overlook full 3D structure or textual context, ShapeWords generates diverse yet consistent images that reflect both the target shape's geometry and the textual description. Experimental results show that ShapeWords produces images that are more text-compliant, aesthetically plausible, while also maintaining 3D shape awareness.

文生图3D引导图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。