用细粒度情绪原型生成有沉浸感的图像,让情感更精准可控。
EmoSpace: Immersive Affective Image Generation Guided by Fine-Grained Emotion Prototypes
- 构建256个可适配的细粒度情绪原型库,实现情绪的精准表达
- 通过原型引导生成,使图像情感对齐度提升且画质保持高水准
- 适合沉浸式场景设计、故事创作等需情绪精准控制的应用
沉浸式情感内容生成旨在创造具有可控情感细微差别的视觉吸引力虚拟现实图像,但现有方法多依赖粗略标签或仅靠提示词控制。尽管现代扩散变换器(DiTs)如FLUX提升了视觉保真度,却未设计用于整合结构化情感表征。我们提出EmoSpace,一个由细粒度情绪原型引导的沉浸式情感内容生成框架,通过三个协同组件将自由文本与情绪描述转化为细粒度情感图像。首先,为表达超越传统分类模型的子情绪差异,EmoSpace通过视觉-语言对齐学习包含256个原型的分层库,并支持输入条件下的自适应调整。其次,原型条件引导通过多路径注入与时间融合,将原型转化为符合DiT的生成信号;迭代提示优化则引入与原型对齐的子情绪描述以丰富提示。第三,情感基础调制协调情感条件与可控制的LoRAs,实现全景化、风格化及多条件生成。定量与定性评估表明,EmoSpace在保持高美学质量的同时显著提升细粒度情感对齐。用户研究显示,相比基线结果,EmoSpace生成内容被认为更具情感一致性,更适用于沉浸式场景设计。此外,我们发现沉浸式呈现会改变情绪感知并增强情感投入。这些发现为情感感知型生成系统在沉浸式媒体中的设计提供依据,潜在应用包括教育、沉浸式叙事和艺术创作。我们将开源代码与模型以促进后续研究。
原文摘要 · Abstract (English)
Immersive affective content generation aims to create visually compelling VR imagery with controllable emotional nuance, yet existing methods typically rely on coarse labels or prompt-only control. Although modern diffusion transformers (DiTs) such as FLUX improve visual fidelity, they are not designed to incorporate structured affective representations. We present EmoSpace, an immersive affective content generation framework guided by fine-grained emotion prototypes, transforming free-form text and emotion descriptions into fine-grained affective imagery through three coordinated components. First, to represent sub-emotion variation beyond conventional categorical models, EmoSpace learns a hierarchical bank of 256 prototypes with input-conditioned adaptation through vision-language alignment. Second, Prototype-Conditioned Steering converts these prototypes into DiT-compatible generation signals through multi-pathway injection and temporal blending, while Iterative Prompt Refinement enriches prompts with prototype-aligned sub-emotion descriptors. Third, Affect-Grounded Modulation coordinates emotion conditioning with controllable LoRAs for panoramic, stylized, and multi-conditional generation. Through quantitative and qualitative evaluations, EmoSpace improves fine-grained emotional alignment while maintaining high aesthetic quality. Our user study shows that EmoSpace outputs are perceived as more emotionally aligned than baseline results and more suitable for immersive scene design. Additionally, we find that immersive presentation alters emotional perception and increases emotional engagement. Together, these findings inform the design of emotion-aware generative systems for immersive media, with potential applications including education, immersive storytelling, and artistic creation. We will release our code and models to facilitate future research along this line.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。