用音频中的电网频率区分录音来源,准确率达92%。
InterGridNet: An Electric Network Frequency Approach for Audio Source Location Classification Using Convolutional Neural Networks
- 基于频谱分析分组,用残差块提取音频特征
- 在SP Cup 2016数据集上达到92%分类准确率
- 适合对音频地理溯源感兴趣的从业者
提出一种名为InterGridNet的新框架,利用浅层RawNet模型对SP Cup 2016数据集中的电力网络频率(ENF)签名进行地理位置分类。数据预处理阶段,根据音频固有特性分为音频与电力组,并通过光谱分析进一步划分为50 Hz和60 Hz组。分类模型中的残差块提取帧级嵌入,通过Softmax激活实现决策。采用神经架构搜索优化浅层RawNet的拓扑结构与超参数。在测试录音中,InterGridNet整体准确率达92%,优于所测试的现有方法。结果表明,该方法能有效区分来自不同电力网格的音频记录,推动了地理定位估计技术的发展。
原文摘要 · Abstract (English)
A novel framework, called InterGridNet, is introduced, leveraging a shallow RawNet model for geolocation classification of Electric Network Frequency (ENF) signatures in the SP Cup 2016 dataset. During data preparation, recordings are sorted into audio and power groups based on inherent characteristics, further divided into 50 Hz and 60 Hz groups via spectrogram analysis. Residual blocks within the classification model extract frame-level embeddings, aiding decision-making through softmax activation. The topology and the hyperparameters of the shallow RawNet are optimized using a Neural Architecture Search. The overall accuracy of InterGridNet in the test recordings is 92%, indicating its effectiveness against the state-of-the-art methods tested in the SP Cup 2016. These findings underscore InterGridNet's effectiveness in accurately classifying audio recordings from diverse power grids, advancing state-of-the-art geolocation estimation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。