arXiv:2410.05320eess.AScs.AI2024-10中稿 · "2024 29th IEEE Sy…被引 2

用简单模型实现高精度语音分类,支持分布式部署。

The OCON model: an old but gold solution for distributable supervised classification

  • 基于单类学习与单网络架构,设计轻量级分类模型。
  • 在语音音素分类任务中达到90.0%-93.7%准确率。
  • 适合需要可扩展性和泛化能力的语音识别场景。

本文将单类学习方法与单类一网络模型系统性地应用于监督分类任务,聚焦自动语音识别领域的元音音素分类案例研究。通过伪神经架构搜索与超参数调优实验,采用有指导的网格搜索策略,模型性能达到与当前复杂架构相当的水平(准确率90.0%-93.7%)。尽管结构简单,该模型更注重语言上下文的泛化能力与分布式应用可行性,获得相关统计与性能指标支持。实验代码已公开于GitHub。

原文摘要 · Abstract (English)

This paper introduces to a structured application of the One-Class approach and the One-Class-One-Network model for supervised classification tasks, specifically addressing a vowel phonemes classification case study within the Automatic Speech Recognition research field. Through pseudo-Neural Architecture Search and Hyper-Parameters Tuning experiments conducted with an informed grid-search methodology, we achieve classification accuracy comparable to nowadays complex architectures (90.0 - 93.7%). Despite its simplicity, our model prioritizes generalization of language context and distributed applicability, supported by relevant statistical and performance metrics. The experiments code is openly available at our GitHub.

语音识别单类学习分布式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。