arXiv:2511.19982cs.CVcs.AI2025-11被引 4

用视觉语言模型反馈优化连续情绪图像生成,提升情感连贯性与准确性。

EmoFeedback$^2$: Reinforcement of Continuous Emotional Image Generation via LVLM-based Reward and Textual Feedback

  • 利用微调的视觉语言模型提供情感奖励与文本反馈,实现生成-理解-反馈闭环。
  • 在自建数据集和公开数据集上均超越现有最优方法,显著提升情感连续性与保真度。
  • 适合关注情绪可控图像生成、多轮交互式创作的研究者与开发者。

连续情绪图像内容生成(C-EICG)因其能同时匹配用户描述与连续情绪值而迅速发展。然而,现有方法缺乏对生成图像的情感反馈,限制了情绪连续性的控制;且简单的文情绪对齐无法根据图像内容自适应调整提示,导致情感保真度不足。为此,我们提出一种新型生成-理解-反馈强化范式(EmoFeedback²),利用微调的大视觉语言模型(LVLM)提供奖励与文本反馈,以生成高质量连续情绪图像。具体地,设计了情感感知奖励反馈策略,由LVLM评估生成图像的情感值并计算与目标情感的奖励,指导生成模型的强化微调,增强图像情感连续性。此外,构建自促进文本反馈框架,使LVLM迭代分析生成图像的情感内容,并自适应生成下一轮提示的优化建议,以细粒度提升情感保真度。大量实验表明,该方法在自建数据集与公开数据集上均有效生成符合预期情绪的高质量图像,优于现有最先进方法。

原文摘要 · Abstract (English)

Continuous emotional image content generation (C-EICG) is emerging rapidly due to its ability to produce images aligned with both user descriptions and continuous emotional values. However, existing approaches lack emotional feedback from generated images, limiting the control of emotional continuity. Additionally, their simple emotion-text alignment fails to adaptively adjust emotional prompts according to image content, leading to insufficient emotional fidelity. To address these concerns, we propose a novel generation-understanding-feedback reinforcement paradigm (EmoFeedback$^2$) for C-EICG, which exploits the reasoning capability of the fine-tuned large vision-language model (LVLM) to provide reward and textual feedback for generating high-quality images with continuous emotions. Specifically, we introduce an emotion-aware reward feedback strategy, where the LVLM evaluates the emotional values of generated images and computes the reward against target emotions, guiding the reinforcement fine-tuning of the generative model and enhancing the emotional continuity of images. Furthermore, we design a self-promotion textual feedback framework, in which the LVLM iteratively analyzes the emotional content of generated images and adaptively produces refinement suggestions for the next-round prompt, improving the emotional fidelity with fine-grained content. Extensive experimental results demonstrate that our approach effectively generates high-quality images with the desired emotions, outperforming existing state-of-the-art methods on both our custom dataset and public dataset.

情绪生成视觉语言模型强化学习图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。