用储层计算直接处理原始音频,无需手工特征提取。
Direct Raw Audio Signal Processing via Reservoir Computing: An Investigation into 'Feature-Free' Architectures

- 采用储层计算架构,直接从原始波形分类音频信号。
- 并行储层结构在多个数据集上表现最优,且模型复杂度低。
- 适合资源受限场景下的低功耗语音识别系统部署。
本文评估了储层计算(Reservoir Computing, RC)作为音频处理的自主、'无特征'框架的潜力,旨在消除传统信号处理中依赖人工设计的特征提取环节。研究探讨高维时序动态的储层能否作为端到端处理器,直接对原始声学信号进行分类。通过跳过诸如梅尔频率倒谱系数(MFCCs)等计算密集型表示,该方法试图缓解传统信号流水线中的重大认知与预处理瓶颈。我们比较了浅层、序列型及并行深层储层架构,以评估其分层特征表示能力。实验结果表明,所提出的并行架构在性能上持续优于浅层和序列基线,同时保持低模型复杂度。这些发现凸显了储层计算在时域音频处理中的高效性与可扩展性,为实现低功耗、预处理极少的可部署声学系统提供了有前景的路径。
原文摘要 · Abstract (English)
This paper evaluates Reservoir Computing (RC) as an autonomous, 'feature-free' framework for audio processing, designed to eliminate traditional, handcrafted feature extraction stages. We investigate whether the high-dimensional temporal dynamics inherent in a reservoir can function as a robust end-to-end processor for the direct classification of raw acoustic signals. By bypassing computationally intensive representations like MFCCs, this approach seeks to mitigate significant intellectual and pre-processing bottlenecks in traditional signal pipelines. Our study evaluates and compares shallow, sequential, and parallel deep reservoir architectures to determine their capacity for hierarchical feature representation. Experimental results demonstrate that the proposed parallel approach consistently outperforms shallow and sequential baselines while maintaining low model complexity. These findings highlight the potential of RC as an efficient and scalable alternative for time-domain audio processing, offering a promising pathway toward deployable, low-power acoustic systems with minimal preprocessing requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。