arXiv:2509.21375cs.CVcs.AI2025-09

自动生成反常识提示,让AI画出不合常理的创意图像

Automated Prompt Generation for Creative and Counterfactual Text-to-image Synthesis

  • 用三模块框架自动改写提示词,实现反常识图像生成
  • 构建首个反常识尺寸图文数据集,性能提升114%
  • 适合创意设计、艺术生成等需要突破常规的场景

文本到图像生成虽因大规模多模态训练取得显著进展,但细粒度控制仍是关键挑战。反常识可控性——即有意识生成违背常识模式的图像——虽具挑战性,却对激发创造力和探索性应用至关重要。本文聚焦反常识尺寸(如在巨按钮旁画一只小海豹),提出一种自动提示工程框架,将基础提示改写为适用于反常识图像的优化提示。框架包含三个组件:图像评估器用于指导数据集构建,监督式提示重写器生成新提示,以及经DPO训练的排序器选择最优提示。我们构建了首个反常识尺寸文本-图像数据集,并通过改进Grounded SAM提升图像评估器性能,相较基线提升114%。实验表明,该方法超越现有最先进基线及ChatGPT-4o,为未来反常识可控性研究奠定基础。

原文摘要 · Abstract (English)

Text-to-image generation has advanced rapidly with large-scale multimodal training, yet fine-grained controllability remains a critical challenge. Counterfactual controllability, defined as the capacity to deliberately generate images that contradict common-sense patterns, remains a major challenge but plays a crucial role in enabling creativity and exploratory applications. In this work, we address this gap with a focus on counterfactual size (e.g., generating a tiny walrus beside a giant button) and propose an automatic prompt engineering framework that adapts base prompts into revised prompts for counterfactual images. The framework comprises three components: an image evaluator that guides dataset construction by identifying successful image generations, a supervised prompt rewriter that produces revised prompts, and a DPO-trained ranker that selects the optimal revised prompt. We construct the first counterfactual size text-image dataset and enhance the image evaluator by extending Grounded SAM with refinements, achieving a 114 percent improvement over its backbone. Experiments demonstrate that our method outperforms state-of-the-art baselines and ChatGPT-4o, establishing a foundation for future research on counterfactual controllability.

文本生成反常识提示工程创意图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。