arXiv:2510.09926cs.LGcs.AI2025-10

用复数神经网络保留音频相位信息,提升分类效果

Phase-Aware Deep Learning with Complex-Valued CNNs for Audio Signal Applications

  • 设计复数卷积网络,完整保留并利用音频相位特征
  • 在音乐流派分类中,引入相位信息使准确率明显提升
  • 适合做音频处理、语音识别等需要相位信息的任务

本研究探索复数卷积神经网络(CVCNNs)在音频信号处理中的设计与应用,重点在于保留常被实值网络忽略的相位信息。提出复数卷积、池化、Wirtinger微分及多种复数激活函数,并改进训练技术如复数批归一化和权重初始化以保证训练稳定。实验分三阶段:首先在标准图像数据集上验证性能,与实值CNN相当,且对合成复数扰动鲁棒;其次在梅尔频率倒谱系数(MFCCs)上进行音频分类,实值输入下复数网络略优;最后引入图神经网络(GNN)通过边权重建模相位,在二分类与多分类音乐流派任务中均取得可测量性能提升。结果表明复数架构具备强表达能力,相位是音频处理中可挖掘的有效特征。尽管现有方法已显潜力,尤其在使用cardioid激活函数时,未来相位感知结构的设计仍需深化。

原文摘要 · Abstract (English)

This study explores the design and application of Complex-Valued Convolutional Neural Networks (CVCNNs) in audio signal processing, with a focus on preserving and utilizing phase information often neglected in real-valued networks. We begin by presenting the foundational theoretical concepts of CVCNNs, including complex convolutions, pooling layers, Wirtinger-based differentiation, and various complex-valued activation functions. These are complemented by critical adaptations of training techniques, including complex batch normalization and weight initialization schemes, to ensure stability in training dynamics. Empirical evaluations are conducted across three stages. First, CVCNNs are benchmarked on standard image datasets, where they demonstrate competitive performance with real-valued CNNs, even under synthetic complex perturbations. Although our focus is audio signal processing, we first evaluate CVCNNs on image datasets to establish baseline performance and validate training stability before applying them to audio tasks. In the second experiment, we focus on audio classification using Mel-Frequency Cepstral Coefficients (MFCCs). CVCNNs trained on real-valued MFCCs slightly outperform real CNNs, while preserving phase in input workflows highlights challenges in exploiting phase without architectural modifications. Finally, a third experiment introduces GNNs to model phase information via edge weighting, where the inclusion of phase yields measurable gains in both binary and multi-class genre classification. These results underscore the expressive capacity of complex-valued architectures and confirm phase as a meaningful and exploitable feature in audio processing applications. While current methods show promise, especially with activations like cardioid, future advances in phase-aware design will be essential to leverage the potential of complex representations in neural networks.

音频处理复数网络相位信息分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。