arXiv:2508.03535cs.CV2025-08被引 8

用心理启发的结构生成情绪明确、可扩展的图像,解决抽象情感表达难题。

CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation

  • 用多模态大模型生成聚焦情绪触发内容的高质量描述,增强语义指导
  • 提出分层低秩适配模块,协同建模情绪共性特征与个性语义
  • 构建大规模情绪艺术数据集EmoArt,支持情绪驱动的艺术创作

情感图像内容生成(EICG)旨在基于给定情绪类别生成语义清晰且情感忠实的图像,具有广泛的应用前景。尽管近期文本到图像扩散模型在生成具体概念方面表现优异,但在处理抽象情绪时仍存在困难。现有专门针对EICG的方法过度依赖词级属性标签进行引导,导致语义不连贯、模糊且可扩展性差。为此,我们提出CoEmoGen,一个以语义一致性与高可扩展性为特点的新流程。具体而言,借助多模态大语言模型(MLLMs),我们构建聚焦情绪触发内容的高质量图像描述,提供上下文丰富的语义引导;同时,受心理学启发,设计分层低秩适配(HiLoRA)模块,协同建模极性共享的低层特征与情绪特异的高层语义。大量实验表明,CoEmoGen在情感忠实度与语义一致性上均优于现有方法,涵盖定量评估、定性分析和用户研究。为直观展示可扩展性,我们构建了大型情绪艺术图像数据集EmoArt,为情绪驱动的艺术创作提供无限灵感。代码与数据集已开源:https://github.com/yuankaishen2001/CoEmoGen。

原文摘要 · Abstract (English)

Emotional Image Content Generation (EICG) aims to generate semantically clear and emotionally faithful images based on given emotion categories, with broad application prospects. While recent text-to-image diffusion models excel at generating concrete concepts, they struggle with the complexity of abstract emotions. There have also emerged methods specifically designed for EICG, but they excessively rely on word-level attribute labels for guidance, which suffer from semantic incoherence, ambiguity, and limited scalability. To address these challenges, we propose CoEmoGen, a novel pipeline notable for its semantic coherence and high scalability. Specifically, leveraging multimodal large language models (MLLMs), we construct high-quality captions focused on emotion-triggering content for context-rich semantic guidance. Furthermore, inspired by psychological insights, we design a Hierarchical Low-Rank Adaptation (HiLoRA) module to cohesively model both polarity-shared low-level features and emotion-specific high-level semantics. Extensive experiments demonstrate CoEmoGen's superiority in emotional faithfulness and semantic coherence from quantitative, qualitative, and user study perspectives. To intuitively showcase scalability, we curate EmoArt, a large-scale dataset of emotionally evocative artistic images, providing endless inspiration for emotion-driven artistic creation. The dataset and code are available at https://github.com/yuankaishen2001/CoEmoGen.

图像生成情绪识别多模态扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。