用情绪维度优化提示词,让AI画画更传神。
EPIG: Emotion-Based Prompting for Personalised Image Generation

- 基于情绪维度构建提示词,不改模型也能增强情感表达。
- 在10个提示词上降低平均唤醒度误差14%~17%,效果显著。
- 适合资源有限或个性化生成,尤其对人物动物类主题更有效。
文本到图像的扩散模型在生成高质量图像方面取得了显著进展,但现有提示策略仍较通用,难以准确传达情感意图和细微情绪特征。本文提出EPIG方法,在生成前通过心理认知启发的情绪表示(效价-唤醒度)与结构化角色感知提示增强,提升提示词的情感相关成分,无需修改或重新训练图像生成主干网络。增强后的情绪感知提示引导生成更具情感一致性的视觉输出,尤其在控制唤醒度方面表现突出。EPIG轻量、无需训练,适用于资源受限及个性化图像生成场景。在包含10个不同提示词的基准测试中,相比强基线(直接插入和LLM提示扩展),EPIG分别将平均唤醒度误差降低14%和12%,差异具有统计显著性。同时保持效价一致性与语义连贯性,如CLIPScore评分所示,且在含人类、儿童或动物等显性主体的提示中,误差降低达17%,体现方法对主体敏感性。
原文摘要 · Abstract (English)
Text-to-image diffusion models have achieved impressive results in synthesizing high-quality images from natural language prompts. However, commonly used prompting strategies remain relatively generic, limiting the model's ability to accurately express emotional intent and nuanced affective attributes. This work proposes EPIG, a method that enhances emotional expressiveness at the prompt level prior to image generation. Grounded in psychologically informed emotion representations (valence-arousal) and leveraging structured, role-aware prompt enrichment, EPIG enriches emotion-related components of prompts without modifying or retraining the image generation backbone. The resulting emotion-aware prompts guide the generative process toward more emotionally coherent visual outputs, with particular effectiveness in controlling arousal. EPIG is lightweight, training-free, and well suited for resource-constrained and personalized image generation scenarios. Experimental results on a benchmark of 10 diverse prompts show that EPIG reduces mean arousal error compared to strong baselines, including naive insertion and LLM-based prompt expansion, with reductions of 14% and 12%, respectively. These improvements are statistically significant. EPIG also preserves valence alignment and semantic consistency, as measured by CLIPScore and supported by ablation studies. The effect is more pronounced on prompts containing explicit subjects such as humans, children, or animals, where the reduction reaches 17%, highlighting the subject-sensitive behavior of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。