将设计评价指标转化为强化学习奖励,提升文本生成图像的图形设计质量。
GDB-Reward: From Evaluation Metrics to Training Rewards for Graphic Design

- 把多种设计评价指标统一为强化学习奖励信号
- 在不微调模型的前提下显著提升排版、色彩和布局准确性
- 适合需要精确设计约束的AI绘图应用
文本到图像模型在自然图像生成上表现优异,但在图形设计中因需满足字体、布局、色彩和视觉传达等精确约束而表现不佳。尽管提示优化可替代昂贵的扩散模型微调,但对冻结的图像生成器进行提示学习仍需有信息量的奖励函数,而生成过程完全不可导。强化学习不要求可导目标,只需能对候选输出排序的标量奖励。这引发一个简单问题:设计评估指标能否直接作为强化学习奖励?我们提出GDB-Reward框架,系统地将异构的图形设计评估指标转化为统一的强化学习奖励。实验表明,GDB-Reward提供了有效的优化目标,在感知质量、渲染保真度和空间准确性方面显著提升设计规范遵循度,同时保持图像生成器完全冻结。更广泛而言,我们的结果表明,异构且不可导的评估指标可从被动基准转向有效优化目标,适用于缺乏可导监督的领域。
原文摘要 · Abstract (English)
Text-to-image models excel at natural image synthesis but struggle with graphic design, where success depends on satisfying precise constraints on typography, layout, color, and visual communication. While prompt optimization offers an attractive alternative to expensive diffusion model fine-tuning, learning prompts for frozen image generators requires informative reward functions despite the entirely non-differentiable generation process. Reinforcement learning does not require differentiable objectives; it requires only scalar rewards capable of ranking candidate outputs. This raises a simple question: can design evaluation metrics themselves become reinforcement learning rewards? Our central contribution is GDB-Reward, a framework that systematically transforms heterogeneous graphic design evaluation metrics into a unified reinforcement learning reward. Experiments demonstrate that GDB-Reward provides an effective optimization objective, substantially improving adherence to the design specification in perceptual quality, rendering fidelity, and spatial accuracy while keeping the image generator entirely frozen. More broadly, our results demonstrate that heterogeneous, non-differentiable evaluation metrics can move beyond passive benchmarking to become effective optimization objectives for reinforcement learning in domains where differentiable supervision is unavailable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。