arXiv:2603.00918cs.CVcs.AI2026-03中稿 · CVPR被引 1

用模型自身生成结果的可信度来优化图像生成质量。

Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards

  • 通过重噪声自生成图像并测量还原精度,评估模型自信度。
  • 在组合生成和图文对齐上显著提升,且无需外部标注数据。
  • 适合追求高质量图文生成而无标注资源的研究者使用。

文本到图像生成推动了设计、媒体和数据增强等领域的内容创作。后训练文本到图像生成模型是提升人类偏好一致性、事实性和美学表现的有前景路径。我们提出SOLACE(自我生成潜在置信度估计),一种后训练框架,将外部奖励监督替换为内部自信心信号:对模型自身的输出进行重噪声处理,并衡量其恢复注入噪声的准确度,低重建误差视为高自信心。SOLACE将此内在信号转化为强化学习的标量奖励,无需外部奖励模型、标注者或偏好数据。通过强化高置信度生成,SOLACE在组合生成、文字渲染和图文对齐方面实现持续提升。与外部奖励结合使用可产生互补改进,并缓解奖励欺骗问题。

原文摘要 · Abstract (English)

Text-to-image generation powers content creation across design, media, and data augmentation. Post-training of text-to-image generative models is a promising path to improve human preference alignment, factuality, and aesthetics. We introduce SOLACE (Self-Originating LAtent Confidence Estimation), a post-training framework that replaces external reward supervision with an internal self-confidence signal: we re-noise the model's own outputs and measure how accurately it recovers the injected noise, treating low reconstruction error as high self-confidence. SOLACE converts this intrinsic signal into scalar rewards for reinforcement learning, requiring no external reward models, annotators, or preference data. By reinforcing high-confidence generations, SOLACE delivers consistent gains in compositional generation, text rendering, and text-image alignment. Integrating SOLACE with external rewards yields complementary improvements while alleviating reward hacking.

图像生成自信心强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。