arXiv:2505.18699cs.CV2025-05IJCV被引 11

用文字控制图片情绪,让图像表达特定情感

Affective Image Editing: Shaping Emotional Factors via Text Descriptions

  • 构建连续情绪谱,将抽象情绪描述转为可操作的视觉语义
  • 通过多模态大模型监督训练,确保编辑后图像触发指定情绪
  • 适合需要情感化图像生成的设计师、内容创作者

日常生活中,图像作为常见的情绪刺激源应用广泛。尽管文本驱动图像编辑已取得显著进展,但针对用户情绪需求的研究仍有限。本文提出AIEdiT,一种基于文本描述的情感图像编辑方法,通过自适应调整图像全局多个情绪因子来激发特定情绪。为表征通用情绪先验,我们构建连续情绪谱并提取细腻情绪请求;为操控情绪因子,设计情绪映射器将视觉抽象的情绪描述转化为具象语义表示;为保证编辑结果激发特定情绪,引入多模态大模型(MLLM)监督训练。推理时,策略性扭曲视觉元素并重塑对应情绪因子,以响应用户指令。此外,我们构建了一个大规模数据集,包含对齐情绪的图文对,用于训练与评估。大量实验表明,AIEdiT性能优越,能有效反映用户的内在情绪诉求。

原文摘要 · Abstract (English)

In daily life, images as common affective stimuli have widespread applications. Despite significant progress in text-driven image editing, there is limited work focusing on understanding users' emotional requests. In this paper, we introduce AIEdiT for Affective Image Editing using Text descriptions, which evokes specific emotions by adaptively shaping multiple emotional factors across the entire images. To represent universal emotional priors, we build the continuous emotional spectrum and extract nuanced emotional requests. To manipulate emotional factors, we design the emotional mapper to translate visually-abstract emotional requests to visually-concrete semantic representations. To ensure that editing results evoke specific emotions, we introduce an MLLM to supervise the model training. During inference, we strategically distort visual elements and subsequently shape corresponding emotional factors to edit images according to users' instructions. Additionally, we introduce a large-scale dataset that includes the emotion-aligned text and image pair set for training and evaluation. Extensive experiments demonstrate that AIEdiT achieves superior performance, effectively reflecting users' emotional requests.

情感编辑文本生成多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。