无需训练即可精准注入情绪,保持图像结构不变
EmoKGEdit: Training-free Affective Injection via Visual Cue Transformation
- 构建多模态情感关联知识图谱,分离情绪与内容特征
- 在不改变图像布局的前提下,有效注入目标情绪
- 适合需要高保真情绪编辑的视觉生成场景
现有图像情绪编辑方法难以将情绪线索与潜在内容表示解耦,常导致情绪表达弱且视觉结构失真。为此,我们提出 EmoKGEdit,一种全新的无训练图像情绪编辑框架,实现精确且结构保持的编辑。我们构建多模态情感关联知识图谱(MSA-KG),解耦物体、场景、属性、视觉线索与情绪之间的复杂关系,显式编码物体-属性-情绪的因果链,作为外部知识引导多模态大模型推理出合理的相关情绪视觉线索,并生成连贯指令。此外,基于 MSA-KG,设计解耦的结构-情绪编辑模块,在潜在空间中明确分离情绪属性与布局特征,确保目标情绪有效注入的同时严格保持视觉空间一致性。大量实验表明,EmoKGEdit 在情绪保真度与内容保留方面均表现优异,优于当前最先进方法。
原文摘要 · Abstract (English)
Existing image emotion editing methods struggle to disentangle emotional cues from latent content representations, often yielding weak emotional expression and distorted visual structures. To bridge this gap, we propose EmoKGEdit, a novel training-free framework for precise and structure-preserving image emotion editing. Specifically, we construct a Multimodal Sentiment Association Knowledge Graph (MSA-KG) to disentangle the intricate relationships among objects, scenes, attributes, visual clues and emotion. MSA-KG explicitly encode the causal chain among object-attribute-emotion, and as external knowledge to support chain of thought reasoning, guiding the multimodal large model to infer plausible emotion-related visual cues and generate coherent instructions. In addition, based on MSA-KG, we design a disentangled structure-emotion editing module that explicitly separates emotional attributes from layout features within the latent space, which ensures that the target emotion is effectively injected while strictly maintaining visual spatial coherence. Extensive experiments demonstrate that EmoKGEdit achieves excellent performance in both emotion fidelity and content preservation, and outperforms the state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。