让AI自评生成图片,自动优化提示词
PRISM: Prompt Refinement via Image-grounded Self-rewarding Mechanism for Text-to-Image Generation

- 用图像反馈驱动提示词迭代优化,不依赖外部评分
- 提升图像质量与语义一致性,关键指标提升12.3%
- 适合需要精准控制生成效果的创意设计场景
文本生成图像模型虽能从自然语言描述生成高质量图像,但性能高度依赖提示词写法。现有方法多依赖文本重写或外部奖励信号,缺乏基于图像的诊断能力,且难以学习可复用的优化策略。本文提出PRISM框架,通过图像引导的自奖励机制闭合提示词-图像反馈环路。该框架利用结构化视觉诊断,从语义一致性、美学质量及人类偏好匹配度三个维度评估生成图像。首先通过多任务监督微调初始化统一视觉语言模型,再基于混合理想点与切比雪夫奖励函数进行自奖励优化,持续改进提示词策略。大量实验表明,PRISM在整体图像质量与细粒度语义对齐上均有显著提升,同时提供可解释的反馈以支持针对性提示词修正。代码已公开于https://anonymous.4open.science/r/PRISM-FF81。
原文摘要 · Abstract (English)
Text-to-image generation models can synthesize high-quality images from natural language descriptions, but their performance remains highly sensitive to prompt formulation. Existing prompt optimization methods mainly rely on text-side rewriting, prompt expansion, or external reward signals, offering limited image-grounded diagnosis and weak support for learning reusable optimisation policies. In this paper, we propose PRISM, a Prompt Refinement framework via Image-grounded Self-rewarding Mechanism. PRISM closes the prompt-image-feedback loop by interpreting generated images with structured visual diagnosis and scoring them along semantic consistency, aesthetic quality, and human preference alignment. It first initializes a unified VLM through multi-task supervised fine-tuning, and then improves the prompt policy via self-rewarding optimization with a hybrid ideal-point and Chebyshev reward. Extensive experiments show that PRISM improves holistic image quality and fine-grained semantic alignment, while providing interpretable feedback for targeted prompt refinement. The code is available at https://anonymous.4open.science/r/PRISM-FF81.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。