arXiv:2508.18569cs.CLcs.CV2025-08被引 3

用结构化提示和轻量强化学习提升视觉隐喻生成的语义一致性。

The Mind's Eye: A Multi-Faceted Reward Framework for Guiding Visual Metaphor Generation

  • 提出S-T-M三元映射提示法,拆解隐喻源、目标与意义。
  • 训练免费方案在分解、CLIP和语义对齐得分上超越GPT-4o与Imagen。
  • 适合追求低算力下高质量抽象隐喻生成的研究者与开发者。

视觉隐喻生成旨在根据输入文本隐喻生成图像,需理解语言并保持概念间的语义连贯性。本文提出一种自评估框架,结合现有指标与新提出的隐喻分解评分及语义对齐(MA)度量。设计两种新方法:无训练流水线通过显式分解提示为源-目标-意义(S-T-M)三元组进行图像合成;互补的训练型流水线则基于自评估奖励机制改进对齐,无需大规模重训。在测试集上,无训练方法在分解、CLIP与MA得分上超越强基线(GPT-4o、Imagen),训练型方法紧随其后。用户研究显示,尽管整体偏好GPT-4o,但无训练方案优于开源模型,并在抽象隐喻上接近甚至超越Imagen。分析表明,S-T-M提示对长或抽象隐喻更有效,闭源模型在短而具体的案例中占优;同时对采样设置敏感。总体而言,结构化提示与轻量强化学习可在低算力下实现良好对齐,当前与人类偏好差距主要源于美学与采样策略。

原文摘要 · Abstract (English)

Visual metaphor generation is a challenging task that aims to generate an image given an input text metaphor. Inherently, it needs language understanding to bind a source concept with a target concept, in a way that preserves meaning while ensuring visual coherence. We propose a self-evaluating visual metaphor generation framework that focuses on metaphor alignment. Our self-evaluation approach combines existing metrics with our newly proposed metaphor decomposition score and a meaning alignment (MA) metric. Within this setup, we explore two novel approaches: a training-free pipeline that explicitly decomposes prompts into source-target-meaning (S-T-M) mapping for image synthesis, and a complementary training-based pipeline that improves alignment using our proposed self-evaluation reward schema, without any large-scale retraining. On the held-out test set, the training-free approach surpasses strong closed baselines (GPT-4o, Imagen) on decomposition, CLIP, and MA scores, with the training-based approach close behind. We evaluate our framework output using a user-facing study, and observed that participants preferred GPT-4o overall, while our training-free pipeline led open-source methods and edged Imagen on abstract metaphors. Our analyses show S-T-M prompting helps longer or more abstract metaphors, with closed models excelling on short, concrete cases; we also observe sensitivity to sampler settings. Overall, structured prompting and lightweight RL perform metaphor alignment well under modest compute, and remaining gaps to human preference appear driven by aesthetics and sampling.

视觉隐喻提示工程生成质量轻量训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。