arXiv:2602.23003eess.SPcs.AI2026-02

用散射变换提升听觉注意力解码,尤其在KUL数据集上表现显著。

Scattering Transform for Auditory Attention Decoding

  • 用两层散射变换替代传统预处理方法,提取更丰富的听觉特征。
  • 在KUL数据集上,对未知说话人分类准确率显著提升。
  • 适合神经网络模型与脑电数据结合的研究者参考。

由于人口结构变化,助听器使用量将上升。新一代助听器需解决鸡尾酒会问题,一种可能方案是基于脑电图的听觉注意力解码。近年研究多采用相同预处理方法。本文提出用散射变换作为替代方案,对比了两层散射变换、常规滤波器组、同步挤压短时傅里叶变换及通用预处理方法。在比利时鲁汶大学(KUL)和丹麦技术大学(DTU)两个常用数据集上,评估了多种分类任务,涵盖卷积神经网络、循环神经网络及最新的Transformer/图神经网络模型。重点考察训练中未见说话人的分类性能。结果表明,两层散射变换在个体相关条件下显著提升性能,尤其在KUL数据集上效果明显;而在DTU数据集上,仅部分模型或在10折交叉验证等更大训练数据下有效。这说明散射变换能提取额外有用信息。

原文摘要 · Abstract (English)

The use of hearing aids will increase in the coming years due to demographic change. One open problem that remains to be solved by a new generation of hearing aids is the cocktail party problem. A possible solution is electroencephalography-based auditory attention decoding. This has been the subject of several studies in recent years, which have in common that they use the same preprocessing methods in most cases. In this work, in order to achieve an advantage, the use of a scattering transform is proposed as an alternative to these preprocessing methods. The two-layer scattering transform is compared with a regular filterbank, the synchrosqueezing short-time Fourier transform and the common preprocessing. To demonstrate the performance, the known and the proposed preprocessing methods are compared for different classification tasks on two widely used datasets, provided by the KU Leuven (KUL) and the Technical University of Denmark (DTU). Both established and new neural-network-based models, CNNs, LSTMs, and recent Transformer/graph-based models are used for classification. Various evaluation strategies were compared, with a focus on the task of classifying speakers who are unknown from the training. We show that the two-layer scattering transform can significantly improve the performance for subject-related conditions, especially on the KUL dataset. However, on the DTU dataset, this only applies to some of the models, or when larger amounts of training data are provided, as in 10-fold cross-validation. This suggests that the scattering transform is capable of extracting additional relevant information.

脑电解码听觉注意散射变换语音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。