用低频谱分析语音节奏,提升印地语系语言识别准确率
Exploring rhythm formant analysis for Indic language classification
- 基于低频谱的节奏形式特征分析,提取语音节奏模式
- 组合特征达69.21%准确率,优于单一R-formants特征
- 适合语音节奏研究与南亚语言识别任务
本文首次对五种印度语言(孟加拉语、卡纳达语、马拉雅拉姆语、马拉地语和泰米尔语)的定量频域节奏线索进行研究。采用由Gibbon提出的节奏形式特征(R-formants)分析方法,通过幅度调制和频率调制包络的低频谱分析来刻画语音节奏。计算了包括R-formants、基于离散余弦变换的度量及频谱度量在内的多种特征。结果显示,阈值法与频谱特征优于直接计算的R-formants;从低频谱图中提取的时间节奏模式具有更强的语言区分能力。综合所有特征后,在五类语言分类任务中达到69.21%的准确率和69.18%的加权F1分数。本研究展示了RFA在印地语系语言节奏表征中的潜力。
原文摘要 · Abstract (English)
This paper reports a preliminary study on quantitative frequency domain rhythm cues for classifying five Indian languages: Bengali, Kannada, Malayalam, Marathi, and Tamil. We employ rhythm formant (R-formants) analysis, a technique introduced by Gibbon that utilizes low-frequency spectral analysis of amplitude modulation and frequency modulation envelopes to characterize speech rhythm. Various measures are computed from the LF spectrum, including R-formants, discrete cosine transform-based measures, and spectral measures. Results show that threshold-based and spectral features outperform directly computed R-formants. Temporal pattern of rhythm derived from LF spectrograms provides better language-discriminating cues. Combining all derived features we achieve an accuracy of 69.21% and a weighted F1 score of 69.18% in classifying the five languages. This study demonstrates the potential of RFA in characterizing speech rhythm for Indian language classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。