arXiv:2609.07409cs.AIcs.LG2026-09

轻量级情感识别框架,实时监控更高效

RAFM-SER++: A Lightweight Multimodal Emotion Recognition Framework for Real-Time Behavioral Monitoring in Surveillance Systems

论文配图:RAFM-SER++: A Lightweight Multimodal Emotion Recognition Framework for Real-Time Behavioral Monitoring in Surveillance Systems
图 1 · 摘自论文原文
  • 采用单向残差注意力融合,降低计算开销
  • 参数减少60%以上,推理速度达79.60帧/秒
  • 适合资源受限的实时安防场景使用

当前多模态语音情感识别系统依赖计算密集的跨模态变换器,难以部署于低延迟、资源受限的监控系统。为此,我们提出RAFM-SER++,一种基于非对称残差注意力融合机制的轻量级框架。该机制通过单向残差注意力路径将情感语音线索注入语义文本表示,避免双向交互带来的高开销。结合自监督跨模态对齐目标与注意力引导池化,有效提升多模态表征学习能力,同时保持低计算负担。在IEMOCAP和ESD数据集上的实验表明,RAFM-SER++显著优于HuBERT-Base基线,在性能上超越SOTA方法MemoCMT:参数减少超过60%,推理速度达79.60 it/s,IEMOCAP上取得81.10%的平衡准确率,ESD上达95.39%。结果证明,轻量级非对称多模态融合是实时监控应用的有效替代方案。

原文摘要 · Abstract (English)

Recent multimodal Speech Emotion Recognition (SER) systems achieve high accuracy through interaction-heavy cross-modal transformers, but their computational cost limits deployment in latency-sensitive and resource-constrained surveillance systems. To address this challenge, we propose RAFM_SER++, a lightweight multimodal SER framework featuring an asymmetric Residual Attention Fusion Mechanism (RAFM). Rather than relying on computationally expensive bidirectional interactions, RAFM injects affective speech cues into semantic text representations through a one-directional residual attention pathway. Combined with a BYOL-inspired cross-modal alignment objective and attention-guided pooling, the proposed framework improves multimodal representation learning while maintaining low computational overhead. Experiments on the IEMOCAP and ESD benchmarks demonstrate that RAFM_SER++ consistently outperforms the HuBERT-Base baseline and achieves a superior accuracy-efficiency trade-off compared with the state-of-the-art MemoCMT. Specifically, RAFM_SER++ reduces trainable parameters by more than 60%, achieves faster inference (79.60 it/s), and attains BACC scores of 81.10% on IEMOCAP and 95.39% on ESD. These results indicate that lightweight asymmetric multimodal fusion is an effective alternative to interaction-heavy cross-modal transformers for real-time surveillance applications.

情感识别轻量化多模态实时监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。