arXiv:2502.14110cs.SDeess.AS2025-02被引 2

用频谱可视图分析声纹,识别准确率高。

On the application of Visibility Graphs in the Spectral Domain for Speaker Recognition

  • 将语音频谱转为可视图,提取拓扑特征
  • 基于决策树集成模型实现高精度声纹识别
  • 适合研究语音特征表示与生物识别的学者

本研究探索了频谱域中可视图在说话人识别中的应用。成年参与者录制了五个西班牙语元音。针对每个发音,基于语音产生源-滤波器模型计算频谱,其中共振峰由声道作为具有共振频率的被动滤波器塑造。频谱轮廓表现出显著的个体内部一致性,反映个体声道解剖结构差异,而跨说话人存在明显变化。从这些频谱轮廓构建可视图,并提取多种图论度量以捕捉其拓扑特征。这些度量组合成代表每位说话人五个元音的特征向量。利用基于这些特征的决策树集成模型,实现了高精度的说话人识别。分析识别出关键拓扑特征对区分说话人至关重要。研究表明,可视图在频谱分析中有效,具有应用于真实语音识别系统的潜力。本研究通过利用频谱域语音信号的拓扑特性,扩展了说话人识别的特征提取工具箱。

原文摘要 · Abstract (English)

In this study, we explore the potential of visibility graphs in the spectral domain for speaker recognition. Adult participants were instructed to record vocalizations of the five Spanish vowels. For each vocalization, we computed the frequency spectrum considering the source-filter model of speech production, where formants are shaped by the vocal tract acting as a passive filter with resonant frequencies. Spectral profiles exhibited consistent intra-speaker characteristics, reflecting individual vocal tract anatomies, while showing variation between speakers. We then constructed visibility graphs from these spectral profiles and extracted various graph-theoretic metrics to capture their topological features. These metrics were assembled into feature vectors representing the five vowels for each speaker. Using an ensemble of decision trees trained on these features, we achieved high accuracy in speaker identification. Our analysis identified key topological features that were critical in distinguishing between speakers. This study demonstrates the effectiveness of visibility graphs for spectral analysis and their potential in speaker recognition. We also discuss the robustness of this approach, offering insights into its applicability for real-world speaker recognition systems. This research contributes to expanding the feature extraction toolbox for speaker recognition by leveraging the topological properties of speech signals in the spectral domain.

声纹识别频谱分析可视图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。