arXiv:2509.15261eess.AScs.SD2025-09中稿 · APSIPA ASC 2025

用预训练自编码器提升声光转换的抗噪能力,适合边缘设备部署。

Pre-training Autoencoder for Acoustic Event Classification via Blinky

  • 用预训练自编码器提取音频紧凑特征,增强声光转换鲁棒性
  • 在15Hz带宽下,宏平均F1优于传统方法
  • 专为树莓派4等边缘设备设计,内存占用低

在基于Blinky的声学事件分类框架中,音频信号被转化为LED光信号,并由单个视频摄像头捕获。然而,30 fps的光学传输通道仅能传递正常音频带宽的约0.2%,且极易受噪声干扰。本文提出一种新型声光转换方法,利用预训练自编码器(AE)的编码器从录制音频中提炼出紧凑且具有判别性的特征。为预训练该自编码器,采用抗噪学习策略:在训练过程中向编码器的潜在表示注入人工噪声,从而提升模型对信道噪声的鲁棒性。编码器架构专门针对当代边缘设备(如树莓派4)的内存限制进行设计。在ESC-50数据集上的模拟实验中,于严格的15 Hz带宽约束下,所提方法的宏平均F1得分高于传统声光转换方法。

原文摘要 · Abstract (English)

In the acoustic event classification (AEC) framework that employs Blinkies, audio signals are converted into LED light emissions and subsequently captured by a single video camera. However, the 30 fps optical transmission channel conveys only about 0.2% of the normal audio bandwidth and is highly susceptible to noise. We propose a novel sound-to-light conversion method that leverages the encoder of a pre-trained autoencoder (AE) to distill compact, discriminative features from the recorded audio. To pre-train the AE, we adopt a noise-robust learning strategy in which artificial noise is injected into the encoder's latent representations during training, thereby enhancing the model's robustness against channel noise. The encoder architecture is specifically designed for the memory footprint of contemporary edge devices such as the Raspberry Pi 4. In a simulation experiment on the ESC-50 dataset under a stringent 15 Hz bandwidth constraint, the proposed method achieved higher macro-F1 scores than conventional sound-to-light conversion approaches.

声学事件分类边缘计算自编码器声光转换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。