arXiv:2412.10426cs.CVcs.CL2024-12ICCV被引 8

为广告图像生成设计创意、对齐与说服力评估指标

CAP: Evaluation of Persuasive and Creative Image Generation

  • 提出创意、提示对齐、说服力三维度评估框架
  • 发现现有模型对隐含提示缺乏有效响应能力
  • 提供简单方法提升图像生成的创意与说服力

我们针对广告图像生成任务,提出三个评估指标:创意(Creativity)、提示对齐(prompt Alignment)和说服力(Persuasiveness),统称为CAP。尽管文本到图像(T2I)生成技术在生成明确描述的高质量图像方面取得进展,但其评估仍面临挑战。现有方法主要关注显式详细描述的对齐性,而对视觉隐含提示的对齐性评估仍是开放问题。此外,创意与说服力是增强广告图像效果的关键品质,却很少被量化。为此,我们提出三项新指标以评估生成图像的创意、对齐与说服力。研究发现,当前T2I模型在面对隐含提示时,在创意、说服力及对齐性方面均表现不佳。我们进一步提出一种简单有效的改进方法,显著提升T2I模型生成更契合、更具创意且更具说服力图像的能力。

原文摘要 · Abstract (English)

We address the task of advertisement image generation and introduce three evaluation metrics to assess Creativity, prompt Alignment, and Persuasiveness (CAP) in generated advertisement images. Despite recent advancements in Text-to-Image (T2I) generation and their performance in generating high-quality images for explicit descriptions, evaluating these models remains challenging. Existing evaluation methods focus largely on assessing alignment with explicit, detailed descriptions, but evaluating alignment with visually implicit prompts remains an open problem. Additionally, creativity and persuasiveness are essential qualities that enhance the effectiveness of advertisement images, yet are seldom measured. To address this, we propose three novel metrics for evaluating the creativity, alignment, and persuasiveness of generated images. Our findings reveal that current T2I models struggle with creativity, persuasiveness, and alignment when the input text is implicit messages. We further introduce a simple yet effective approach to enhance T2I models' capabilities in producing images that are better aligned, more creative, and more persuasive.

图像生成广告设计评估指标T2I

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。