将文本情绪转化为具象图像,让情感表达更直观
Towards Deeper Emotional Reflection: Crafting Affective Image Filters with Generative Priors
- 用多模态模型从文本生成带情绪的视觉图像
- 新模型在情感准确性和内容一致性上优于现有方法
- 适合社交媒体情感表达与创意设计场景
社交平台允许用户通过图文结合表达情绪。本文提出情感图像滤镜(AIF)任务,旨在将文本中的抽象情绪转化为具象视觉图像,生成更具情感冲击力的结果。我们首先构建了AIF数据集并定义模型形式,提出基于多模态Transformer的AIF-B作为初步尝试。随后,提出AIF-D进一步深化情感表达,有效利用预训练大规模扩散模型的生成先验。定量与定性实验表明,AIF模型在内容一致性和情感忠实度方面均优于当前最优方法。大量用户研究显示,AIF模型更能有效激发特定情绪。基于结果,我们全面讨论了AIF模型的价值与潜力。
原文摘要 · Abstract (English)
Social media platforms enable users to express emotions by posting text with accompanying images. In this paper, we propose the Affective Image Filter (AIF) task, which aims to reflect visually-abstract emotions from text into visually-concrete images, thereby creating emotionally compelling results. We first introduce the AIF dataset and the formulation of the AIF models. Then, we present AIF-B as an initial attempt based on a multi-modal transformer architecture. After that, we propose AIF-D as an extension of AIF-B towards deeper emotional reflection, effectively leveraging generative priors from pre-trained large-scale diffusion models. Quantitative and qualitative experiments demonstrate that AIF models achieve superior performance for both content consistency and emotional fidelity compared to state-of-the-art methods. Extensive user study experiments demonstrate that AIF models are significantly more effective at evoking specific emotions. Based on the presented results, we comprehensively discuss the value and potential of AIF models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。