用文本控制图片情绪,实现精细情感迁移。
EmoLat: Text-driven Image Sentiment Transfer via Emotion Latent Space
- 构建情绪潜在空间,关联文本语义与视觉情绪特征。
- 在自建数据集上优于现有方法,情感迁移更准确。
- 适合需要可控图像编辑的研究者与开发者。
我们提出EmoLat,一种新型情绪潜在空间,通过建模文本语义与视觉情绪特征间的跨模态关联,实现细粒度、文本驱动的图像情感迁移。在EmoLat中,构建情绪语义图以捕捉情绪、物体与视觉属性之间的关系结构;为增强情绪表征的可区分性与迁移能力,采用对抗正则化对齐多模态间的潜在情绪分布。基于EmoLat,提出跨模态情感迁移框架,通过联合嵌入文本与EmoLat特征来操控图像情感。网络使用包含语义一致性、情绪对齐与对抗正则化的多目标损失进行优化。为支持有效建模,构建了大规模基准数据集EmoSpace Set,包含图像在情绪、物体语义与视觉属性上的密集标注。在EmoSpace Set上的大量实验表明,本方法在定量指标与定性迁移保真度上均显著优于现有最先进方法,确立了文本引导可控图像情感编辑的新范式。EmoSpace Set与全部代码已公开于http://github.com/JingVIPLab/EmoLat。
原文摘要 · Abstract (English)
We propose EmoLat, a novel emotion latent space that enables fine-grained, text-driven image sentiment transfer by modeling cross-modal correlations between textual semantics and visual emotion features. Within EmoLat, an emotion semantic graph is constructed to capture the relational structure among emotions, objects, and visual attributes. To enhance the discriminability and transferability of emotion representations, we employ adversarial regularization, aligning the latent emotion distributions across modalities. Building upon EmoLat, a cross-modal sentiment transfer framework is proposed to manipulate image sentiment via joint embedding of text and EmoLat features. The network is optimized using a multi-objective loss incorporating semantic consistency, emotion alignment, and adversarial regularization. To support effective modeling, we construct EmoSpace Set, a large-scale benchmark dataset comprising images with dense annotations on emotions, object semantics, and visual attributes. Extensive experiments on EmoSpace Set demonstrate that our approach significantly outperforms existing state-of-the-art methods in both quantitative metrics and qualitative transfer fidelity, establishing a new paradigm for controllable image sentiment editing guided by textual input. The EmoSpace Set and all the code are available at http://github.com/JingVIPLab/EmoLat.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。