arXiv:2502.07556cs.HCcs.CV2025-02被引 25

用草图辅助生成图像,让非专业用户也能轻松控制物体位置和关系。

SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches

  • 通过草图区域自动推断语义提示,减少用户输入负担。
  • 将粗糙草图转为边缘锚点,提升生成图像的空间一致性。
  • 适合希望精准控制多物体布局的设计师或初学者使用。

文本生成图像模型虽能产出视觉吸引人的图像,但非专业用户在撰写合适提示词和指定精细空间条件(如深度或Canny参考)时仍面临挑战,尤其当涉及多个物体时。为此,我们提出SketchFlex,一个基于区域草图的交互式系统,以增强空间条件图像生成的灵活性。该系统利用众包获取的物体属性与关系构建语义空间,自动推断出合理描述的用户提示;同时将用户的粗略草图优化为基于Canny的形状锚点,确保生成质量与用户意图一致。实验表明,SketchFlex在生成图像的语义连贯性上优于端到端模型,显著降低认知负荷,并在匹配用户意图方面优于基于区域的生成基线。

原文摘要 · Abstract (English)

Text-to-image models can generate visually appealing images from text descriptions. Efforts have been devoted to improving model controls with prompt tuning and spatial conditioning. However, our formative study highlights the challenges for non-expert users in crafting appropriate prompts and specifying fine-grained spatial conditions (e.g., depth or canny references) to generate semantically cohesive images, especially when multiple objects are involved. In response, we introduce SketchFlex, an interactive system designed to improve the flexibility of spatially conditioned image generation using rough region sketches. The system automatically infers user prompts with rational descriptions within a semantic space enriched by crowd-sourced object attributes and relationships. Additionally, SketchFlex refines users' rough sketches into canny-based shape anchors, ensuring the generation quality and alignment of user intentions. Experimental results demonstrate that SketchFlex achieves more cohesive image generations than end-to-end models, meanwhile significantly reducing cognitive load and better matching user intentions compared to region-based generation baseline.

文本生成图像草图控制空间一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。