用连续情绪值高效精准编辑图片情感,支持自然场景与社交场景。
MooD: Perception-Enhanced Efficient Affective Image Editing via Continuous Valence-Arousal Modeling

- 直接使用连续情绪维度(愉悦-唤醒)作为编辑指令
- 在多个数据集上实现高保真度与强情感控制力,效率优于现有方法
- 构建新数据集AffectSet,覆盖自然与社交场景,支持更广应用
情感图像编辑(AIE)旨在通过修改视觉内容引发特定情绪。尽管当前方法在编辑质量上表现优异,但普遍忽视推理效率,限制了其在计算社交场景中的应用。此外,多数方法依赖离散情绪表示,难以建模复杂人类情感,制约了交互场景下的表达能力。为此,我们提出MooD,首个直接以连续愉悦-唤醒(Valence-Arousal, VA)值为编辑指令的高效情感图像编辑框架,实现计算社交系统中细粒度且高效的AIE。首先,提出VA感知检索策略,连接模糊情绪值与详细视觉语义;在此基础上,MooD融合视觉迁移与感知增强的语义引导机制,实现可控编辑。考虑到现有VA标注数据集主要聚焦社交场景,严重忽略自然场景,我们构建了AffectSet——一个覆盖多样化场景的综合性VA标注数据集,支持模型优化与评估。大量定性与定量实验表明,MooD在情感可控制性与视觉保真度方面均达到领先水平,同时保持高效率。消融实验证实了设计关键因素的有效性。
原文摘要 · Abstract (English)
Affective Image Editing (AIE) aims to modify visual content to evoke targeted emotions. Although current approaches achieve impressive editing quality, they often overlook inference efficiency, which limits their applicability in computational social scenarios. Moreover, most methods depend on discrete emotion representations, which hinder the continuous modeling of complex human emotions and constrain expressive capabilities in interactive scenarios. To tackle these gaps, we propose MooD, the first framework that directly leverages continuous Valence-Arousal (VA) values as editing instruction for fine-grained and efficient AIE in computational social systems. Specifically, we first introduce a VA-Aware retrieval strategy to bridge vague affective values and detailed visual semantics. Building upon this, MooD integrates visual transfer and perception-enhanced semantic guidance to achieve controllable AIE. Furthermore, considering that existing VA-annotated datasets mainly focus on social scenarios and largely overlook natural scenes, we therefore construct AffectSet, a comprehensive VA-annotated dataset covering diverse scenarios, to support model optimization and evaluation. Extensive qualitative and quantitative experimental results demonstrate that our MooD achieves superior performance in both affective controllability and visual fidelity while maintaining high efficiency. A series of ablation studies further reveal the crucial factors of our design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。