MUSE统一生成与编辑图像情绪,无需训练新模型
MUSE: Manipulating Unified Framework for Synthesizing Emotions in Images via Test-Time Optimization
- 用梯度优化情感令牌,通过现成分类器稳定引导情绪生成
- 根据语义相似度判断最佳引导时机,提升情绪表达准确性
- 多情绪损失减少干扰,适合心理治疗、叙事创作等场景
图像激发的情绪深刻影响感知,常被置于内容之上。现有图像情绪合成(IES)方法将生成与编辑任务人为分离,造成效率低下且限制了二者天然融合的应用场景,如心理干预或故事创作。本文提出MUSE,首个能同时实现情绪生成与编辑的统一框架。受大语言模型和扩散模型领域广泛使用的测试时扩展(Test-Time Scaling, TTS)启发,MUSE无需额外更新扩散模型或专用情绪合成数据集。具体解决三个关键问题:(1) 如何通过梯度优化情感令牌,利用现成情绪分类器稳定引导合成;(2) 在何时引入情绪引导,通过语义相似度作为监督信号识别最优时机;(3) 哪种情绪进行引导,采用多情绪损失降低固有及相似情绪的干扰。实验表明,MUSE在生成与编辑任务中均优于所有现有方法,显著提升情绪准确率与语义多样性,同时保持内容一致性、提示遵循性与真实情绪表达之间的良好平衡,确立了情绪合成的新范式。
原文摘要 · Abstract (English)
Images evoke emotions that profoundly influence perception, often prioritized over content. Current Image Emotional Synthesis (IES) approaches artificially separate generation and editing tasks, creating inefficiencies and limiting applications where these tasks naturally intertwine, such as therapeutic interventions or storytelling. In this work, we introduce MUSE, the first unified framework capable of both emotional generation and editing. By adopting a strategy conceptually aligned with Test-Time Scaling (TTS) that widely used in both LLM and diffusion model communities, it avoids the requirement for additional updating diffusion model and specialized emotional synthesis datasets. More specifically, MUSE addresses three key questions in emotional synthesis: (1) HOW to stably guide synthesis by leveraging an off-the-shelf emotion classifier with gradient-based optimization of emotional tokens; (2) WHEN to introduce emotional guidance by identifying the optimal timing using semantic similarity as a supervisory signal; and (3) WHICH emotion to guide synthesis through a multi-emotion loss that reduces interference from inherent and similar emotions. Experimental results show that MUSE performs favorably against all methods for both generation and editing, improving emotional accuracy and semantic diversity while maintaining an optimal balance between desired content, adherence to text prompts, and realistic emotional expression. It establishes a new paradigm for emotion synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。