arXiv:2501.05710cs.CV2025-01ICCV被引 26

用情绪维度生成带情感的图像,更精准控制画面情绪。

EmotiCrafter: Text-to-Emotional-Image Generation based on Valence-Arousal Model

  • 用唤醒度-效价模型替代离散情绪分类,实现连续情绪控制。
  • 结合文本提示与情绪参数,生成符合指定情感的图像内容。
  • 在情感表达和内容一致性上优于现有方法,适合创意设计应用。

研究表明情绪能增强认知并影响信息传播。尽管视觉情绪分析研究丰富,但帮助用户生成富有情感的图像内容的工作仍有限。现有情感图像生成多依赖离散情绪类别,难以准确捕捉复杂细微的情绪变化,且难以根据文本提示精确控制生成内容。本文提出连续情感图像内容生成(C-EICG)新任务,并构建基于效价-唤醒度模型的EmotiCrafter模型,通过文本提示与效价-唤醒度值生成图像。我们设计了新型情绪嵌入映射网络,将效价-唤醒度值融入文本特征,实现与输入提示一致的特定情绪表达;同时引入新损失函数以强化情绪表现。实验表明,该方法能有效生成具有特定情绪且内容匹配的图像,性能优于现有技术。

原文摘要 · Abstract (English)

Recent research shows that emotions can enhance users' cognition and influence information communication. While research on visual emotion analysis is extensive, limited work has been done on helping users generate emotionally rich image content. Existing work on emotional image generation relies on discrete emotion categories, making it challenging to capture complex and subtle emotional nuances accurately. Additionally, these methods struggle to control the specific content of generated images based on text prompts. In this work, we introduce the new task of continuous emotional image content generation (C-EICG) and present EmotiCrafter, an emotional image generation model that generates images based on text prompts and Valence-Arousal values. Specifically, we propose a novel emotion-embedding mapping network that embeds Valence-Arousal values into textual features, enabling the capture of specific emotions in alignment with intended input prompts. Additionally, we introduce a loss function to enhance emotion expression. The experimental results show that our method effectively generates images representing specific emotions with the desired content and outperforms existing techniques.

图像生成情感计算文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。