arXiv:2607.29679cs.CV2026-07被引 1

发现文本结构化程度影响图像生成质量,可提升生成效果。

Scaling Properties of Text Conditioning in Visual Generation

论文配图:Scaling Properties of Text Conditioning in Visual Generation
图 1 · 摘自论文原文
  • 用语言结构度量方法分析提示词对生成的影响
  • 语言结构越强,扩散损失越低,生成更稳定
  • 适合想提升图像生成精度的开发者和研究者

我们研究了视觉生成中文本条件化的经验缩放规律。由于扩散损失不随自然语言提示中的词元数量变化,这类性质很少被测量。令人惊讶的是,我们发现收敛后的扩散损失与提示中结构化语言量呈正相关。为此,我们采用两种互补度量:白盒似然指标(GPG)和黑盒属性指标(ED)。在控制训练实验中,收敛扩散损失随GPG近似线性下降,随ED呈幂律关系。基于这些缩放规律,我们通过从图像中提取语义与几何标注构建结构化提示,提升了模型的可扩散性;并通过监督微调、冷启动和验证器门控在线策略蒸馏训练提示生成器,增强了提示适应性。该系统在几乎所有组合性、推理和常识基准测试中优于所有评估的开源权重模型,且在多数评测中达到或超过最强闭源模型表现。

原文摘要 · Abstract (English)

We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Surprisingly, we find that the converged diffusion loss scales with the amount of structured language in the prompt. To quantify structured language, we adapt two complementary measures: a white-box likelihood metric (GPG) and a black-box attribute metric (ED). Across controlled training runs, the converged diffusion loss decreases approximately linearly with GPG and follows a power law with ED. Guided by these scaling properties, we improve \emph{diffusability} by constructing structured prompts with semantic and geometric annotations derived from images, and improve \emph{promptability} by training a prompter through supervised fine-tuning, cold-start, and verifier-gated on-policy distillation. The resulting system outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world-knowledge benchmark, while matching or surpassing the strongest closed-weight models on most evaluations.

图像生成文本条件提示工程扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。