arXiv:2602.20672cs.CV2026-02被引 1

让AI图像生成支持精确的框选位置、大小和颜色控制。

BBQ-to-Image: Numeric Bounding Box and Qolor Control in Large-Scale Text-to-Image Models

  • 用带数值标注的结构化文本训练模型,实现精准控制。
  • 在框对齐和颜色保真度上优于现有最佳模型。
  • 适合需要精确布局与配色的专业设计场景。

文本到图像模型在真实感和可控性方面快速进步,近期方法利用长而详细的描述性文本支持细粒度生成。然而,仍存在根本性的参数空白:现有模型依赖描述性语言,而专业工作流程则需要对对象位置、尺寸和颜色进行精确的数值控制。本文提出BBQ,一种大规模文本到图像模型,直接在统一的结构化文本框架中以数值边界框和RGB三元组为条件。通过在包含参数标注的描述文本上训练,无需架构修改或推理时优化,即可实现精确的空间与色彩控制。这还支持直观的用户界面,如拖拽对象和拾色器,用精确熟悉的控制替代模糊的迭代提示。在全面评估中,BBQ实现了出色的框对齐效果,并在RGB颜色保真度上优于当前最优基线。更广泛而言,我们的结果支持一种新范式:将用户意图转化为中间结构化语言,由基于流的Transformer作为渲染器处理,自然兼容数值参数。

原文摘要 · Abstract (English)

Text-to-image models have rapidly advanced in realism and controllability, with recent approaches leveraging long, detailed captions to support fine-grained generation. However, a fundamental parametric gap remains: existing models rely on descriptive language, whereas professional workflows require precise numeric control over object location, size, and color. In this work, we introduce BBQ, a large-scale text-to-image model that directly conditions on numeric bounding boxes and RGB triplets within a unified structured-text framework. We obtain precise spatial and chromatic control by training on captions enriched with parametric annotations, without architectural modifications or inference-time optimization. This also enables intuitive user interfaces such as object dragging and color pickers, replacing ambiguous iterative prompting with precise, familiar controls. Across comprehensive evaluations, BBQ achieves strong box alignment and improves RGB color fidelity over state-of-the-art baselines. More broadly, our results support a new paradigm in which user intent is translated into an intermediate structured language, consumed by a flow-based transformer acting as a renderer and naturally accommodating numeric parameters.

图像生成可控生成数值控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。