构建首个面向虚拟现实的可穿戴多模态情绪识别数据集
Introducing WARM-VR: Benchmark Dataset for Multimodal Wearable Affect Recognition in Virtual Reality

- 用可穿戴设备采集31人多种生理信号,配合虚拟现实场景诱发情绪
- 引入嗅觉刺激后,用户负向情绪显著下降,验证了多感官沉浸效果
- 在情绪分类任务中,轻量级Transformer模型表现最优,适合资源受限场景
随着人机交互融入日常生活,机器学习使系统能更好感知并响应用户情绪状态。现有情绪识别数据集多聚焦静态环境,难以适用于虚拟现实(VR)等沉浸式多媒体场景。本文提出WARM-VR,一个公开的多模态数据集,支持基于可穿戴传感器的沉浸式多感官情绪识别。数据来自31名19-37岁参与者,使用腕带(测量血容积脉搏、皮电反应、皮肤温度、三轴加速度)和胸带(记录心电图)采集生理信号。参与者经历诱导压力的算术任务后,进入虚拟海滩放松体验,期间同步呈现视觉、听觉与嗅觉刺激。情绪状态通过标准化自评量表与生理数据分析双重评估。问卷统计分析表明,虚拟现实放松显著降低负向情绪,尤其在嗅觉增强条件下。我们采用主流机器学习算法建立基准:仅用血容积脉搏信号进行效价二分类时,CNN与CNN-Bi-GRU模型平均F1-score达0.63,AUC为0.69;唤醒度分类中,轻量级Transformer模型表现最佳(F1-0: 0.54,F1-1: 0.63),优于循环混合模型;在放松任务中,CNN-Bi-GRU模型达到最高综合性能(平均F1-score 0.64,AUC 0.69)。
原文摘要 · Abstract (English)
With the growing integration of human-computer interaction into everyday life, advances in machine learning have enabled systems to better perceive and respond to users' emotional states. Most existing affect recognition datasets focus on static environments, limiting their applicability to immersive multimedia contexts such as Virtual Reality (VR). In this paper, we introduce WARM-VR, a novel publicly available multimodal dataset designed to support affect recognition in immersive, multisensory environments using wearable sensing instrumentation. Data were collected from 31 participants aged 19-37 using wearable sensors: a wristband measuring Blood Volume Pulse (BVP), EDA, skin Temperature, three-axis Acceleration, and a chest strap recording ECG signals. Participants engaged in immersive VR experiences designed to elicit relaxation through a calming beach environment following stress induction via an arithmetic task. These sessions incorporated synchronized multimedia stimuli: visual, auditory, and olfactory. Affective states were assessed subjectively through validated self-report questionnaires and objectively through the analysis of physiological measurements. Statistical analysis of the questionnaires confirmed that VR relaxation significantly reduced negative affect, particularly with olfactory enhancement. Furthermore, we established a benchmark on the dataset using widely recognized machine learning algorithms. The best performance for binary classification from BVP data of valence, was obtained with a CNN and a CNN-Bi-GRU model, both achieving an average F1-score of 0.63 and an AUC of 0.69. For arousal, a lightweight Transformer architecture provided the most balanced results (F1-0 0.54 and F1-1 0.63), outperforming recurrent hybrids. In the relaxation task, a CNN-Bi-GRU model reached the highest overall performance (average F1-score 0.64, AUC 0.69).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。