用多模态生理信号实现嗅觉诱发情绪的高精度识别
Multimodal Wearable-Based Olfactory-Induced Emotion Recognition in Arousal-Valence Dimensions
- 提出时空频混合融合网络,解决多模态信号非平稳与不同步问题
- 在自建数据集上达到92.40%准确率,优于现有方法5.07%
- 适合关注情绪计算、可穿戴设备与多模态融合的研究者
嗅觉通过直接作用于大脑情感回路,成为一种无侵入且认知负担轻的情绪调节途径,对日常及注意力关键场景中的情感计算具有重要意义。然而,现有嗅觉情绪研究存在两大局限:过度关注效价维度而忽视唤醒度;缺乏同步采集中枢与外周生理反应的多模态数据集。为此,本文基于111名受试者构建大规模多模态嗅觉情绪数据集,其中气味在二维唤醒-效价空间中标注,并同步记录脑电(EEG)、心电(ECG)和光电容积脉搏波(PPG)信号。针对多模态信号存在的非平稳性、时延差异与跨模态异质性挑战,提出时空频混合融合网络(STF-HFNet),包含三个核心模块:频率聚合处理学习自适应频率聚合以建模非平稳动态;互导注意力实现无需同步先验的双向时间对齐校准;混合协同融合结合空间与通道注意力机制,在增强跨模态互补性的同时抑制冗余信息。大量实验表明,STF-HFNet在AMIGOS数据集上达到88.34%的识别准确率,在自建数据集上达92.40%,分别超越SOTA方法8.27%和5.07%。
原文摘要 · Abstract (English)
Olfaction is important for emotion regulation because it acts as a non-intrusive and cognitively lightweight pathway that directly engages the brain s affective circuitry and achieves unobtrusive emotional modulation. This trait is essential for advancing practical affective computing in daily and attention-critical scenarios. However, current olfactory emotion research has two key limitations. First, it overemphasises the valence dimension while neglecting arousal. Second, it lacks multimodal datasets that synchronously capture central and peripheral physiological responses to olfactory stimuli. To address these issues, we construct a large-scale multimodal olfactory emotion dataset based on 111 subjects, in which odors are labeled in the 2D arousal-valence space and electroencephalogram (EEG), electrocardiogram (ECG), and photoplethysmography (PPG) signals synchronously recorded. Nevertheless, multimodal signals present challenges such as non-stationarity, differences in latency, and cross-modal heterogeneity. Thus, we propose a spatiotemporal-frequency hybrid fusion network (STF-HFNet), which integrates three core modules. Frequency aggregation processing learns adaptive frequency aggregation in order to model non-stationary dynamics. Reciprocal guided attention enables reciprocal bidirectional calibration for cross-modal temporal alignment without synchronisation priors. Hybrid collaborative fusion combines spatial and channel attention mechanisms to enhance cross-modal complementarity while suppressing redundant information. Extensive experiments show that STF-HFNet achieves state-of-the-art (SOTA) recognition accuracies of 88.34% on the AMIGOS dataset and 92.40% on our self-constructed dataset, and outperform the SOTA methods by 8.27% and 5.07%, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。