arXiv:2509.16760eess.AScs.SD2025-09

用图模型选关键声景特征,发现情绪唤醒与效价强关联。

Feature Selection via Graph Topology Inference for Soundscape Emotion Recognition

  • 基于线性结构方程建模,构建特征间关系的稀疏图
  • 在Emo-Soundscapes数据集上,识别出影响情绪的关键特征组合
  • 提出新拐点检测法,可评估特征选择的不确定性

声景情绪识别(SER)关注声音感知而非单纯噪声水平,常以唤醒度和效价作为情感描述。本文结合图学习与新型信息准则,构建针对Emo-Soundscapes数据集的特征选择框架。通过定制化线性结构方程模型(SEM),推断特征与两类情绪输出间的稀疏图关系。提出一种广义拐点检测器,同时提供最优稀疏度点估计与置信区间。实验包含关系可视化,结果部分验证已有发现,但揭示唤醒度与效价之间存在强关联,挑战了传统假设。

原文摘要 · Abstract (English)

Research on soundscapes has shifted the focus of environmental acoustics from noise levels to the perception of sounds, incorporating contextual factors. Soundscape emotion recognition (SER) models perception using a set of features, with arousal and valence commonly regarded as sufficient descriptors of affect. In this work, we blend \emph{graph learning} techniques with a novel \emph{information criterion} to develop a feature selection framework for SER. Specifically, we estimate a sparse graph representation of feature relations using linear structural equation models (SEM) tailored to the widely used Emo-Soundscapes dataset. The resulting graph captures the relations between input features and the two emotional outputs. To determine the appropriate level of sparsity, we propose a novel \emph{generalized elbow detector}, which provides both a point estimate and an uncertainty interval. We conduct an extensive evaluation of our methods, including visualizations of the inferred relations. While several of our findings align with previous studies, the graph representation also reveals a strong connection between arousal and valence, challenging common SER assumptions.

声景识别特征选择图学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。