用大模型精准操控图像情绪,让生成图更符合主观情感需求。
Towards LLM-centric Affective Visual Customization via Efficient and Precise Emotion Manipulating
- 以大模型为核心,实现情绪语义的高效转换
- 在新数据集上优于现有方法,情绪操控更精准
- 适合需要情感化视觉生成的研究与应用
以往视觉定制研究主要依赖控制信号(如语言、布局、边缘)与图像的客观对齐,忽视了主观情绪内容,且缺乏通用的情感化视觉定制基础模型。为此,本文提出面向大模型的情感化视觉定制(L-AVC)任务,聚焦通过多模态大模型修改图像的主观情绪。进一步指出,在此任务中,如何高效实现语义层面的情绪转换(即跨情绪语义转换)以及如何精确保留非情绪相关的内容(即外情绪语义保留)尤为关键且具有挑战性。为此,本文提出一种高效精准的情绪操控方法(EPEM),包含两个模块:用于高效对齐编辑前后情绪语义的高效跨情绪转换(EIC)模块,以及用于精确保留非情绪内容的精准外情绪保留(PER)模块。在自建的L-AVC数据集上的全面实验表明,该方法显著优于多个先进基线模型,验证了情绪信息在情感化视觉定制中的重要性,以及EPEM在高效精准操控情绪方面的有效性。
原文摘要 · Abstract (English)
Previous studies on visual customization primarily rely on the objective alignment between various control signals (e.g., language, layout and canny) and the edited images, which largely ignore the subjective emotional contents, and more importantly lack general-purpose foundation models for affective visual customization. With this in mind, this paper proposes an LLM-centric Affective Visual Customization (L-AVC) task, which focuses on generating images within modifying their subjective emotions via Multimodal LLM. Further, this paper contends that how to make the model efficiently align emotion conversion in semantics (named inter-emotion semantic conversion) and how to precisely retain emotion-agnostic contents (named exter-emotion semantic retaining) are rather important and challenging in this L-AVC task. To this end, this paper proposes an Efficient and Precise Emotion Manipulating approach for editing subjective emotions in images. Specifically, an Efficient Inter-emotion Converting (EIC) module is tailored to make the LLM efficiently align emotion conversion in semantics before and after editing, followed by a Precise Exter-emotion Retaining (PER) module to precisely retain the emotion-agnostic contents. Comprehensive experimental evaluations on our constructed L-AVC dataset demonstrate the great advantage of the proposed EPEM approach to the L-AVC task over several state-of-the-art baselines. This justifies the importance of emotion information for L-AVC and the effectiveness of EPEM in efficiently and precisely manipulating such information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。