用一张画布整合多种控制指令,实现精准图像生成
Canvas-to-Image: Compositional Image Generation with Multimodal Controls
- 将文本、参考图、位置、姿态等多模态控制融合成一张复合画布
- 在多人物、姿态约束等挑战性任务中,身份保留与控制遵循性显著提升
- 适合需要精细图像编排的设计师和创意工作者
尽管现代扩散模型在生成高质量多样图像方面表现优异,但在高保真度的组合式与多模态控制方面仍存在困难,尤其是在用户同时提供文本提示、主体参考、空间布局、姿态约束和版面标注时。我们提出Canvas-to-Image,一个统一框架,将这些异构控制信号整合到单一画布界面中,使生成图像能忠实反映用户意图。核心思想是将多种控制信号编码为模型可直接理解的复合画布图像,实现视觉-空间联合推理。我们还构建了一套多任务数据集,并提出多任务画布训练策略,使扩散模型在统一学习范式下联合理解并整合异构控制信号。该联合训练使模型在推理时能跨模态推理,而非依赖特定任务启发式规则,且在多控制场景中具有良好泛化能力。大量实验表明,Canvas-to-Image在多人物组合、姿态控制组合、布局约束生成及多控制生成等挑战性基准上,显著优于现有最优方法,在身份保持与控制遵循性方面均有大幅提升。
原文摘要 · Abstract (English)
While modern diffusion models excel at generating high-quality and diverse images, they still struggle with high-fidelity compositional and multimodal control, particularly when users simultaneously specify text prompts, subject references, spatial arrangements, pose constraints, and layout annotations. We introduce Canvas-to-Image, a unified framework that consolidates these heterogeneous controls into a single canvas interface, enabling users to generate images that faithfully reflect their intent. Our key idea is to encode diverse control signals into a single composite canvas image that the model can directly interpret for integrated visual-spatial reasoning. We further curate a suite of multi-task datasets and propose a Multi-Task Canvas Training strategy that optimizes the diffusion model to jointly understand and integrate heterogeneous controls into text-to-image generation within a unified learning paradigm. This joint training enables Canvas-to-Image to reason across multiple control modalities rather than relying on task-specific heuristics, and it generalizes well to multi-control scenarios during inference. Extensive experiments show that Canvas-to-Image significantly outperforms state-of-the-art methods in identity preservation and control adherence across challenging benchmarks, including multi-person composition, pose-controlled composition, layout-constrained generation, and multi-control generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。