用八维情绪分布精准控制图像情绪变化,保留原图结构
AffectDelta: Beyond Emotion Labels for Image Editing

- 将情绪编辑建模为八维分布的转移过程
- 在24.9万对图像上训练,实现跨类别与同类别情绪变换
- 适合需要精细情绪调控的生成任务,如影视特效或心理研究
情感驱动的图像编辑旨在通过修改源图像中与情绪相关的视觉线索,引发指定目标情绪,同时保持原始场景的整体构图和语义结构一致性。现有场景级编辑器通常以单一情绪类别指定目标,且多从操作级文本指令中学习视觉转换。单一类别会将混合情感终点压缩为单一主导标签,而语言无法精确量化共存情绪的增减或稳定程度。我们提出AffectDelta,一种源感知编辑器,将编辑视为八维情绪分布之间的转移。一个冻结的情绪分布预测器估计源状态,其与目标的有符号差值编码了请求转移的方向与幅度。AffectDelta内部的过渡编码器与源感知扩散主干共同将该信号转化为上下文相关的语义与外观变化。为训练此模型,我们构建了AffectPair-249K,包含248,841对源-目标图像对,带有预测的八维情绪分布,涵盖跨类别与同类别转移。与六种基线对比的实验,结合定量评估与定性比较,证明其在情感契合度与内容保真度上的提升,消融实验验证了设计选择的有效性。代码与数据集将在录用后公开。
原文摘要 · Abstract (English)
Emotion-driven image editing aims to evoke a specified target emotion by modifying emotion-relevant visual cues in a source image, while preserving the overall composition and semantic-structural coherence of the original scene. Existing scene-level editors typically specify the target with a single emotion category and often learn visual transformations from operation-level text instructions. A category collapses a mixed affective endpoint into one dominant label, while language cannot precisely quantify how coexisting emotions should increase, decrease, or remain stable. We introduce AffectDelta, a source-aware editor that treats editing as a transition between eight-dimensional emotion distributions. A frozen Emotion Distribution Predictor estimates the source state, and the signed source-to-target difference encodes the direction and magnitude of the requested transition. Within AffectDelta, an internal transition encoder and a source-aware diffusion backbone jointly translate this signal into context-dependent semantic and appearance changes. To train this formulation, we construct AffectPair-249K, comprising 248,841 source-target pairs with predicted eight-dimensional distributions and spanning both cross-category and within-category transitions. Experiments against six baselines, combining quantitative evaluation with qualitative comparisons, demonstrate improved affective alignment and content preservation, while ablations validate our design choices. Code and dataset will be made publicly available upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。