arXiv:2604.00933cs.CV2026-04被引 1

构建120万张带情感与感知属性的图像数据集,支持精细情绪控制生成。

EmoScene: A Dual-space Dataset for Controllable Affective Image Generation

  • 双空间设计:同时包含情绪标签与连续维度(VAD)及可测量外观属性
  • 1.2万张图像经人工审核,情绪标注一致性达91.22%,κ值0.85
  • 基于该数据集开发的AffectCtrl模型实现85.75%情绪准确率,显著优于现有方法

文本到图像扩散模型虽具备高视觉保真度,但细粒度情感控制仍困难,因文本情感提示常无法精确描述影响情绪表达的视觉感知因素。现有视觉-情感数据集多局限于离散标签、特定领域或感知属性监督不足。本文提出EmoScene,一个大规模双空间数据集,涵盖超过300个场景类别,共120万张图像。其情感空间联合表示离散情绪与连续的效价-唤醒-支配(VAD)维度,感知空间记录可量化的外观属性,上下文描述则将二者与场景语义对齐。EmoScene结合多模型标注与人工质检流程,随机审计30,519张图像后,离散情绪标签一致率达91.22%,独立多评估者评价得Fleiss' κ=0.85。在控制源图像与场景结构后,情感维度与感知属性在不同数据源间保持稳定关联,反映统计趋势而非确定性规则。为验证数据集价值,进一步构建AffectCtrl模型,通过学习冻结扩散模型条件空间中的残差,实现类别情绪生成与连续控制(VAD、亮度、饱和度)。AffectCtrl在共享评估协议下,情绪分类准确率达85.75%,在效价和唤醒控制上优于EmotiCrafter,五维连续控制的皮尔逊相关系数达0.673–0.765。结果表明,EmoScene为视觉生成中情感表达的分析与控制提供了可扩展的数据基础。

原文摘要 · Abstract (English)

Text-to-image diffusion models achieve high visual fidelity, yet fine-grained affective control remains difficult because textual emotion cues often fail to specify the visual perceptual factors underlying affective expression. Existing visual-affect datasets are likewise often limited to discrete labels, specific domains, or limited supervision of perceptual attributes. We introduce EmoScene, a large-scale dual-space dataset for controllable affective image generation, containing 1.2M images across more than 300 scene categories. Its affective space jointly represents discrete emotions and continuous valence--arousal--dominance (VAD), its perceptual space records measurable appearance attributes, and contextual descriptions ground both in scene semantics. EmoScene combines multi-model annotation with human-in-the-loop quality control. A random audit of 30,519 images yields 91.22\% agreement on discrete emotion labels, while an independent multi-rater evaluation yields Fleiss' $κ=0.85$. After controlling for source and scene composition, affective dimensions and perceptual attributes exhibit stable associations across data sources, reflecting statistical tendencies rather than deterministic visual rules. To demonstrate the dataset's utility, we further develop AffectCtrl, which learns residuals in the conditioning space of frozen diffusion models to support categorical emotion generation and continuous control over VAD, brightness, and saturation. AffectCtrl achieves 85.75\% categorical emotion accuracy, outperforms EmotiCrafter in valence and arousal control under a shared evaluation protocol, and obtains Pearson correlations of 0.673--0.765 across all five continuous axes. These results demonstrate that EmoScene provides a scalable data foundation for analyzing and controlling affective expression in visual generation.

情感生成图像数据集扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。