轻量级模型实现城市声学监测高精度语音分类,兼顾效率与泛化。
The OCON model: an old but green solution for distributable supervised classification for acoustic monitoring in smart cities
- 结合单类学习与参数优化,构建简单高效分类框架。
- 在语音识别任务中达到复杂模型相当的准确率,能耗更低。
- 适合资源受限场景,如智能城市中的实时声环境监控。
本文探索了单类方法及单类单网络模型在监督分类任务中的结构化应用,聚焦于语音识别(ASR)领域的元音音素分类与说话人识别。以专有传感与照明系统为载体,该模型用于监测城市街道的声环境与空气污染。通过融合伪神经架构搜索与超参数调优,采用启发式网格搜索方法,实现了与当前复杂模型相媲美的分类准确率,并深入分析了说话人识别与能效表现。尽管模型结构简单,但其在语言和性别泛化方面具有较强潜力,适用于计算资源受限的广泛场景,相关实验代码已开源。
原文摘要 · Abstract (English)
This paper explores a structured application of the One-Class approach and the One-Class-One-Network model for supervised classification tasks, focusing on vowel phonemes classification and speakers recognition for the Automatic Speech Recognition (ASR) domain. For our case-study, the ASR model runs on a proprietary sensing and lightning system, exploited to monitor acoustic and air pollution on urban streets. We formalize combinations of pseudo-Neural Architecture Search and Hyper-Parameters Tuning experiments, using an informed grid-search methodology, to achieve classification accuracy comparable to nowadays most complex architectures, delving into the speaker recognition and energy efficiency aspects. Despite its simplicity, our model proposal has a very good chance to generalize the language and speaker genders context for widespread applicability in computational constrained contexts, proved by relevant statistical and performance metrics. Our experiments code is openly accessible on our GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。