用框和消失线精准控制图像比例与透视,提升生成可控性。
Proportion and Perspective Control for Flow-Based Image Generation
- 用边界框控制物体位置与大小,用消失线调节三维场景构图。
- 在简单约束下控制效果显著,复杂场景仍存局限。
- 适合需要精确布局的艺术创作或设计类应用。
现代文本到图像的扩散模型虽能生成高保真图像,但在空间和几何结构控制方面能力有限。为此,我们提出并评估了两种专用于艺术控制的ControlNet:(1) 比例ControlNet,利用边界框指定物体的位置与尺度;(2) 透视ControlNet,通过消失线控制场景的三维几何结构。我们采用基于视觉-语言模型的数据标注管道和专门的条件图像合成算法来支持这两个模块的训练。实验表明,两者均能有效提供控制,但在复杂约束下仍有不足。两个模型已发布于HuggingFace:https://huggingface.co/obvious-research。
原文摘要 · Abstract (English)
While modern text-to-image diffusion models generate high-fidelity images, they offer limited control over the spatial and geometric structure of the output. To address this, we introduce and evaluate two ControlNets specialized for artistic control: (1) a proportion ControlNet that uses bounding boxes to dictate the position and scale of objects, and (2) a perspective ControlNet that employs vanishing lines to control the 3D geometry of the scene. We support the training of these modules with data pipelines that leverage vision-language models for annotation and specialized algorithms for conditioning image synthesis. Our experiments demonstrate that both modules provide effective control but exhibit limitations with complex constraints. Both models are released on HuggingFace: https://huggingface.co/obvious-research
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。