用纳米忆阻器网络直接处理原始音频,实现低延迟高效率语音分类。
Memristive Nanowire Network for Energy Efficient Audio Classification: Pre-Processing-Free Reservoir Computing with Reduced Latency
- 用忆阻纳米线网络直接从原始音频提取紧凑特征,无需传统预处理。
- 在亚秒级训练内实现98.95%准确率,数据压缩达66倍,比传统方法快255倍。
- 适合边缘设备部署,尤其适用于资源受限、低延迟语音识别场景。
高效音频特征提取对低延迟、资源受限的语音识别至关重要。传统预处理方法如梅尔频谱图、感知线性预测(PLP)和可学习频谱图虽精度高,但需大量特征和计算开销。类脑计算在低延迟与低功耗方面具有显著优势。本文首次将忆阻纳米线网络作为类脑硬件预处理层,用于语音数字分类。该网络可直接从原始音频中提取紧凑且信息丰富的特征,在保持高精度的同时实现显著的数据压缩与训练效率提升。相比最先进软件方法,纳米线特征在子秒级训练下达到98.95%准确率(数据压缩66倍,采用XGBoost)和97.9%准确率(压缩255倍,采用随机森林)。在多种分类器上,纳米线特征始终超过90%准确率且压缩比超62.5倍,优于传统最优特征(如MFCC),在性能不降的前提下大幅提升效率。此外,其在多说话人音频分类中达到96.5%准确率,为当前最高水平,兼具最大压缩比与最低训练时间。纳米线预处理还能增强音频数据的线性可分性,提升简单分类器性能并跨说话人泛化。结果表明,忆阻纳米线网络提供了一种新颖、低延迟、数据高效的特征提取方案,支持高性能类脑音频分类。
原文摘要 · Abstract (English)
Efficient audio feature extraction is critical for low-latency, resource-constrained speech recognition. Conventional preprocessing techniques, such as Mel Spectrogram, Perceptual Linear Prediction (PLP), and Learnable Spectrogram, achieve high classification accuracy but require large feature sets and significant computation. The low-latency and power efficiency benefits of neuromorphic computing offer a strong potential for audio classification. Here, we introduce memristive nanowire networks as a neuromorphic hardware preprocessing layer for spoken-digit classification, a capability not previously demonstrated. Nanowire networks extract compact, informative features directly from raw audio, achieving a favorable trade-off between accuracy, dimensionality reduction from the original audio size (data compression) , and training time efficiency. Compared with state-of-the-art software techniques, nanowire features reach 98.95% accuracy with 66 times data compression (XGBoost) and 97.9% accuracy with 255 times compression (Random Forest) in sub-second training latency. Across multiple classifiers nanowire features consistently achieve more than 90% accuracy with more than 62.5 times compression, outperforming features extracted by conventional state-of-the-art techniques such as MFCC in efficiency without loss of performance. Moreover, nanowire features achieve 96.5% accuracy classifying multispeaker audios, outperforming all state-of-the-art feature accuracies while achieving the highest data compression and lowest training time. Nanowire network preprocessing also enhances linear separability of audio data, improving simple classifier performance and generalizing across speakers. These results demonstrate that memristive nanowire networks provide a novel, low-latency, and data-efficient feature extraction approach, enabling high-performance neuromorphic audio classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。