让图像既忠于描述又表达指定情绪,突破内容与情感的平衡难题。
EmoCtrl: Controllable Emotional Image Content Generation
- 通过文本与视觉双增强模块,将情绪语义转化为视觉线索。
- 在保持内容一致性的同时,情感表达显著优于现有方法。
- 适合创意设计、情绪化内容生成等需要精准情感控制的场景。
图像通过视觉内容和情感基调共同传递意义,影响人类感知。我们提出可控情感图像内容生成(C-EICG),旨在生成既忠实于给定内容描述又能表达目标情绪的图像。现有文本到图像模型虽能保证内容一致,但缺乏情感感知;而情感驱动模型则常以内容失真为代价生成情绪化结果。为解决这一问题,我们构建了一个包含内容、情绪与情感提示的标注数据集,将抽象情绪映射至视觉特征。EmoCtrl引入文本与视觉双重情绪增强模块,通过描述性语义和感知线索丰富情感表达。为对齐人类偏好,我们设计了基于情绪奖励的偏好优化机制。全面实验表明,EmoCtrl在内容忠实度与情感表现力上均优于现有方法。用户研究进一步验证其与人类偏好的高度契合。此外,EmoCtrl在创意应用中表现出良好泛化能力,证明所学情感标记具有鲁棒性与可迁移性。
原文摘要 · Abstract (English)
An image conveys meaning through both its visual content and emotional tone, jointly shaping human perception. We introduce Controllable Emotional Image Content Generation (C-EICG), which aims to generate images that remain faithful to a given content description while expressing a target emotion. Existing text-to-image models ensure content consistency but lack emotional awareness, whereas emotion-driven models generate affective results at the cost of content distortion. To address this gap, we propose EmoCtrl, supported by a dataset annotated with content, emotion, and affective prompts, bridging abstract emotions to visual cues. EmoCtrl incorporates textual and visual emotion enhancement modules that enrich affective expression via descriptive semantics and perceptual cues. To align with human preference, we further introduce an emotion-driven preference optimization with specifically designed emotion reward. Comprehensive experiments demonstrate that EmoCtrl achieves faithful content and expressive emotion control, outperforming existing methods. User studies confirm EmoCtrl's strong alignment with human preference. Moreover, EmoCtrl generalizes well to creative applications, further demonstrating the robustness and adaptability of the learned emotion tokens.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。