用情绪驱动图像编辑,让抽象情感变具体视觉变化。
Moodifier: MLLM-Enhanced Emotion-Driven Image Editing
- 构建800万+带情绪标注的图像数据集MoodArchive
- 提出MoodifyCLIP实现情绪到视觉属性的精准映射
- 无需训练即可基于大模型实现情感化编辑,适合创意设计场景
将情感与视觉内容结合以实现情绪驱动的图像编辑在创意产业中潜力巨大,但因情感的抽象性及其在不同情境下的多样表现,精确操控仍具挑战。本文提出集成方案:首先构建包含800万+图像的MoodArchive数据集,由LLaVA生成并经人工部分验证的情绪层级标注;其次开发基于该数据集微调的MoodifyCLIP模型,实现从抽象情绪到具体视觉属性的转换;最后提出Moodifier,一个无需训练的编辑模型,利用MoodifyCLIP与多模态大语言模型(MLLMs),在保持内容完整性的同时实现精准情绪转化。系统适用于人物表情、时尚设计、珠宝与家居装饰等多个领域,支持创作者快速可视化情绪变化且保留身份与结构特征。大量实验表明,Moodifier在情绪准确率与内容保真度上优于现有方法,提供上下文恰当的编辑效果。项目将在论文接受后公开MoodArchive数据集、MoodifyCLIP模型及Moodifier代码与演示。
原文摘要 · Abstract (English)
Bridging emotions and visual content for emotion-driven image editing holds great potential in creative industries, yet precise manipulation remains challenging due to the abstract nature of emotions and their varied manifestations across different contexts. We tackle this challenge with an integrated approach consisting of three complementary components. First, we introduce MoodArchive, an 8M+ image dataset with detailed hierarchical emotional annotations generated by LLaVA and partially validated by human evaluators. Second, we develop MoodifyCLIP, a vision-language model fine-tuned on MoodArchive to translate abstract emotions into specific visual attributes. Third, we propose Moodifier, a training-free editing model leveraging MoodifyCLIP and multimodal large language models (MLLMs) to enable precise emotional transformations while preserving content integrity. Our system works across diverse domains such as character expressions, fashion design, jewelry, and home décor, enabling creators to quickly visualize emotional variations while preserving identity and structure. Extensive experimental evaluations show that Moodifier outperforms existing methods in both emotional accuracy and content preservation, providing contextually appropriate edits. By linking abstract emotions to concrete visual changes, our solution unlocks new possibilities for emotional content creation in real-world applications. We will release the MoodArchive dataset, MoodifyCLIP model, and make the Moodifier code and demo publicly available upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。