arXiv:2503.05929cs.SDcs.AI2025-03被引 1

将语音特征转为彩色图像,提升说话人识别准确率。

Audio-to-Image Encoding for Improved Voice Characteristic Detection Using Deep Convolutional Neural Networks

  • 三通道图像融合原始音频与声学特征,实现多维信息编码
  • 在双说话人数据上达到98%分类准确率,显著优于传统方法
  • 适合语音识别、生物特征分析等需要高精度身份验证的场景

本文提出一种新型音频到图像的编码框架,将语音的多维度特征整合为单张RGB图像用于说话人识别。绿色通道编码原始音频信号,红色通道嵌入语音信号的统计描述符(包括基频、频谱质心、带宽、滚降频率、过零率、梅尔频率倒谱系数MFCCs、RMS能量、频谱平坦度、频谱对比度、音高色度及谐噪比等关键指标的均值与中位数),蓝色通道则以空间化子帧形式呈现这些特征。基于该复合图像训练的深度卷积神经网络,在两说话人数据集上实现了98%的说话人分类准确率,表明这种多通道融合表示能为语音识别任务提供更具判别力的输入。

原文摘要 · Abstract (English)

This paper introduces a novel audio-to-image encoding framework that integrates multiple dimensions of voice characteristics into a single RGB image for speaker recognition. In this method, the green channel encodes raw audio data, the red channel embeds statistical descriptors of the voice signal (including key metrics such as median and mean values for fundamental frequency, spectral centroid, bandwidth, rolloff, zero-crossing rate, MFCCs, RMS energy, spectral flatness, spectral contrast, chroma, and harmonic-to-noise ratio), and the blue channel comprises subframes representing these features in a spatially organized format. A deep convolutional neural network trained on these composite images achieves 98% accuracy in speaker classification across two speakers, suggesting that this integrated multi-channel representation can provide a more discriminative input for voice recognition tasks.

语音识别图像编码深度学习特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。