arXiv:2505.23509cs.SDcs.LG2025-05

用类脑听觉机制实现高效可解释的音频分类

Spectrotemporal Modulation: Efficient and Interpretable Feature Representation for Classifying Speech, Music, and Environmental Sounds

  • 基于人类听觉皮层机制设计声谱时序调制特征
  • 无需预训练即可达到预训练深度模型性能
  • 适合需可解释性的语音、音乐、环境音识别场景

音频深度神经网络在多种机器听觉任务中表现优异,但其表征通常计算成本高且难以解释,仍有优化空间。本文提出一种基于声谱时序调制(STM)特征的新方法,该方法模拟人类听觉皮层的神经生理表征。所提出的STM模型在未进行任何预训练的情况下,对自然语境下的语音、音乐和环境声音分类性能与预训练音频深度神经网络相当。结果表明,STM是一种高效且可解释的音频分类特征表示,有助于推进机器听觉发展,并为理解语音与听觉科学、开发音频脑机接口及认知计算开辟新路径。

原文摘要 · Abstract (English)

Audio DNNs have demonstrated impressive performance on various machine listening tasks; however, most of their representations are computationally costly and uninterpretable, leaving room for optimization. Here, we propose a novel approach centered on spectrotemporal modulation (STM) features, a signal processing method that mimics the neurophysiological representation in the human auditory cortex. The classification performance of our STM-based model, without any pretraining, is comparable to that of pretrained audio DNNs across diverse naturalistic speech, music, and environmental sounds, which are essential categories for both human cognition and machine perception. These results show that STM is an efficient and interpretable feature representation for audio classification, advancing the development of machine listening and unlocking exciting new possibilities for basic understanding of speech and auditory sciences, as well as developing audio BCI and cognitive computing.

音频分类可解释性类脑计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。