arXiv:2412.13486cs.CVcs.CL2024-12被引 2

无需训练的草图到场景生成方法,实现可控概念艺术创作。

T$^3$-S2S: Training-free Triplet Tuning for Sketch to Scene Synthesis in Controllable Concept Art Generation

  • 通过三模块设计优化控制网络,实现草图到多实例场景的精准生成。
  • 在多实例场景生成中保持提示词对齐,细节与输入草图高度一致。
  • 适合游戏、影视等需要结构化地形布局的可控艺术生成场景。

2D概念艺术生成用于3D场景是计算机图形学中的关键但具挑战性的任务,因创建自然直观环境仍需大量手动设计工作。尽管生成式AI已通过文本到图像合成简化了2D概念设计,但在复杂多实例场景和结构化地形布局方面支持有限。本文提出一种无需训练的三重调优方法(T³-S2S),用于草图到场景生成。该方法通过三个核心模块:提示平衡(Prompt Balance)确保关键词表征并减少关键实例遗漏;特征优先(Characteristic Priority)突出草图特征的前K个通道索引;密集调优(Dense Tuning)细化注意力图中实例相关区域的轮廓细节。借助T³-S2S的可控性,我们还引入双提示集共享策略,生成分层感知的等距视图与地形视角表示。实验表明,该草图到场景工作流能持续生成与提示对齐的多实例2D场景。

原文摘要 · Abstract (English)

2D concept art generation for 3D scenes is a crucial yet challenging task in computer graphics, as creating natural intuitive environments still demands extensive manual effort in concept design. While generative AI has simplified 2D concept design via text-to-image synthesis, it struggles with complex multi-instance scenes and offers limited support for structured terrain layout. In this paper, we propose a Training-free Triplet Tuning for Sketch-to-Scene (T3-S2S) generation after reviewing the entire cross-attention mechanism. This scheme revitalizes the ControlNet model for detailed multi-instance generation via three key modules: Prompt Balance ensures keyword representation and minimizes the risk of missing critical instances; Characteristic Priority emphasizes sketch-based features by highlighting TopK indices in feature channels; and Dense Tuning refines contour details within instance-related regions of the attention map. Leveraging the controllability of T3-S2S, we also introduce a feature-sharing strategy with dual prompt sets to generate layer-aware isometric and terrain-view representations for the terrain layout. Experiments show that our sketch-to-scene workflow consistently produces multi-instance 2D scenes with details aligned with input prompts.

概念艺术草图生成控制生成图像合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。