梳理声谱图特征在音频语音分析中的应用与设计选择。
Spectrogram features for audio and speech analysis
- 系统回顾声谱图的分辨率、跨度与数值表示等设计参数。
- 分析不同设置对分类器性能的影响,揭示特征与模型的适配关系。
- 适合音频处理、语音识别研究者参考,尤其关注前端特征设计。
基于声谱图的表征已主导深度学习音频分析系统,常用于语音分析。其核心优势在于将声音以时间-频率二维信号形式呈现,既具备可解释的物理基础,又可利用为图像处理设计的卷积神经网络等机器学习技术。声谱图由两个维度的分辨率与跨度,以及元素的表示与缩放方式共同定义。研究人员已在众多应用场景中探索了多种配置,不同设置对各类任务表现出偏好。本文综述声谱图表征的使用现状,探讨前端特征选择如何与后端分类器架构适配,以实现最佳性能。
原文摘要 · Abstract (English)
Spectrogram-based representations have grown to dominate the feature space for deep learning audio analysis systems, and are often adopted for speech analysis also. Initially, the primary motivator for spectrogram-based representations was their ability to present sound as a two dimensional signal in the time-frequency plane, which not only provides an interpretable physical basis for analysing sound, but also unlocks the use of a wide range of machine learning techniques such as convolutional neural networks, that had been developed for image processing. A spectrogram is a matrix characterised by the resolution and span of its two dimensions, as well as by the representation and scaling of each element. Many possibilities for these three characteristics have been explored by researchers across numerous application areas, with different settings showing affinity for various tasks. This paper reviews the use of spectrogram-based representations and surveys the state-of-the-art to question how front-end feature representation choice allies with back-end classifier architecture for different tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。