arXiv:2412.13541cs.CVcs.LG2024-12被引 3

通过时空模糊多模态元学习,实现小样本下的精准情绪识别。

Spatio-Temporal Fuzzy-oriented Multi-Modal Meta-Learning for Fine-grained Emotion Recognition

  • 分视图构建元任务,融合时空卷积捕捉数据异质性。
  • 引入广义模糊规则处理情绪复杂与模糊问题,提升鲁棒性。
  • 适用于少样本、跨场景情绪识别,适合医疗与个性化服务。

细粒度情绪识别(FER)在疾病诊断、个性化推荐和多媒体挖掘等领域至关重要。然而现有方法面临三大挑战:(i)依赖大量持续标注数据,因情绪真实场景中复杂且模糊,标注成本高;(ii)无法捕捉随时间变化的情绪模式差异,通常假设采样期内时序相关性一致;(iii)忽视不同场景下情绪信息的空间异质性,即数据分布可能存在偏差或干扰。为此,提出时空模糊导向的多模态元学习框架(ST-F2M)。首先将多模态视频划分为多个视图,每个视图对应某一情绪的一种模态,同情绪的随机多视图构成一个元训练任务。接着使用结合空间与时间卷积的集成模块编码每个任务数据,反映时空异质性。再基于广义模糊规则为每个任务添加模糊语义信息,以应对情绪的复杂性与模糊性。最后通过元递归神经网络学习情绪相关的通用元知识,实现快速稳健的细粒度情绪识别。大量实验表明,ST-F2M在准确率与模型效率上均优于多种先进方法。此外,通过消融实验与深入分析,验证了其有效性。

原文摘要 · Abstract (English)

Fine-grained emotion recognition (FER) plays a vital role in various fields, such as disease diagnosis, personalized recommendations, and multimedia mining. However, existing FER methods face three key challenges in real-world applications: (i) they rely on large amounts of continuously annotated data to ensure accuracy since emotions are complex and ambiguous in reality, which is costly and time-consuming; (ii) they cannot capture the temporal heterogeneity caused by changing emotion patterns, because they usually assume that the temporal correlation within sampling periods is the same; (iii) they do not consider the spatial heterogeneity of different FER scenarios, that is, the distribution of emotion information in different data may have bias or interference. To address these challenges, we propose a Spatio-Temporal Fuzzy-oriented Multi-modal Meta-learning framework (ST-F2M). Specifically, ST-F2M first divides the multi-modal videos into multiple views, and each view corresponds to one modality of one emotion. Multiple randomly selected views for the same emotion form a meta-training task. Next, ST-F2M uses an integrated module with spatial and temporal convolutions to encode the data of each task, reflecting the spatial and temporal heterogeneity. Then it adds fuzzy semantic information to each task based on generalized fuzzy rules, which helps handle the complexity and ambiguity of emotions. Finally, ST-F2M learns emotion-related general meta-knowledge through meta-recurrent neural networks to achieve fast and robust fine-grained emotion recognition. Extensive experiments show that ST-F2M outperforms various state-of-the-art methods in terms of accuracy and model efficiency. In addition, we construct ablation studies and further analysis to explore why ST-F2M performs well.

情绪识别多模态元学习模糊逻辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。