揭示建筑环境视觉表征中情绪维度的组织规律
Organization of Valence and Arousal in Vision-Language Representations of Built Environments: Insights from the EMOIS Dataset
- 用CLIP模型分析建筑图像的情绪表征空间结构
- 情绪值预测准确率达0.865(愉悦度)和0.807(唤醒度)
- 提供可解释的界面,适合人机交互与情感计算研究
建筑环境的视觉感知影响人们日常的情感印象,但其在视觉基础模型中的表征方式仍不明确。为此,我们构建了包含1,544张真实建筑环境图像的EMOIS数据集,每张图像由约120名日本成人通过大规模在线调查标注了情绪愉悦度和唤醒度。基于对比语言-图像预训练(CLIP)表示,我们开展预测与几何分析,发现愉悦度在表征空间中具有更强且更一致的组织性。跨数据集分析显示,与通用情绪图片集OASIS相比,建筑环境在情感组织上存在差异。回归分析表明,在重复内部留出评估中,愉悦度和唤醒度的决定系数均值分别达到0.865和0.807。最后,我们展示了一个基于示例的界面,用于可视化预测情绪值的解释。这些结果有助于理解建筑环境的情感表征,并使EMOIS成为该领域未来情感计算研究的重要资源。
原文摘要 · Abstract (English)
Visual perception of built environments contributes to the affective impressions that people form in everyday life. However, how these impressions are represented within vision foundation models remains largely unexplored. To support the systematic investigation of this subject, we introduce the Emotional Impression of Spaces (EMOIS) dataset, comprising 1,544 real-world built-environment images. Each image is annotated with image-evoked valence and arousal ratings collected from Japanese adults by conducting a large-scale web-based survey, with approximately 120 ratings per image. Using Contrastive Language--Image Pre-training (CLIP) representations, we perform predictive and geometric analyses to systematically investigate how valence and arousal are encoded and organized within the representation space. These analyses reveal that valence exhibited stronger and more coherent organization than arousal. Cross-dataset analyses with the Open Affective Standardized Image Set (OASIS), a benchmark dataset of general affective photographs, reveal differences in affective organization between the two datasets. Regression analyses demonstrate high predictive performance for valence and arousal within EMOIS, with mean coefficients of determination of 0.865 and 0.807, respectively, across repeated internal hold-out evaluations. Finally, we present an example-based interface illustrating how learned representations can support qualitative interpretation of predicted affective values. These findings can help elucidate affective representations of built environments and establish EMOIS as a densely annotated resource for future affective computing research in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。