用几何相关性建模提升自监督语音情绪识别性能
Geometric Second-Order Feature Correlation Learning for Self-Supervised Speech Emotion Recognition

- 引入二阶相关层,通过协方差描述特征间协同模式
- 在ESD和RAVDESS上准确率显著优于传统一阶聚合方法
- 适合关注语音情绪识别与自监督表示学习的研究者
自监督学习(SSL)为语音情绪识别(SER)生成了丰富的上下文表征,但如何将这些表征整合为整体描述符仍是一大瓶颈。传统的一阶聚合隐含假设特征独立,忽略了潜在的黎曼几何结构,丢失了高阶关系所蕴含的表征能力。为此,本文提出一种新型二阶相关(SOC)层,不将特征孤立处理,而是将其相关性建模为协方差描述符,捕捉协同共现模式,作为鲁棒情绪识别的判别性签名。通过对数-欧氏映射(LEM)将这些描述符从黎曼流形投影到欧氏切空间,既保持几何完整性,又支持直接线性判别学习。在ESD和RAVDESS数据集上的大量实验表明,SOC能有效恢复一阶池化中丢失的判别信息,并成功聚合高维SSL特征。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) yields powerful, context-rich representations for speech emotion recognition (SER), yet aggregating these representations into holistic descriptors remains a bottleneck. Conventional first-order aggregation implicitly assumes feature independence, which overlooks the latent Riemannian geometry and discards higher-order relationships essential to the representational power of the backbone. To address this problem, this paper proposes a novel Second-Order Correlation (SOC) layer. Instead of treating features in isolation, SOC models feature correlations as covariance descriptors to capture synergistic co-occurrence patterns, which serve as discriminative signatures for robust emotion recognition. By mapping these descriptors from the Riemannian manifold to a Euclidean tangent space through Log-Euclidean mapping (LEM), the proposed method preserves geometric integrity while enabling direct linear discriminative learning. Extensive experiments on the ESD and RAVDESS datasets demonstrate that SOC recovers discriminative information lost in first-order pooling and effectively aggregates high-dimensional SSL features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。