用频域分析方法区分阿萨姆邦两种低资源语言的语音节奏差异
Analyzing long-term rhythm variations in Mising and Assamese using frequency domain correlates
- 基于幅度与频率调制的低频谱图提取六阶节奏共振峰轨迹
- 结合二维离散余弦变换特征,实现两类语言节奏的准确区分
- 无需语音标注即可研究低资源语言的节奏结构,适合语音学研究者
本文研究了印度东北部阿萨姆邦两种低资源语言——米辛语和阿萨姆语的长期语音节奏变化。通过分析由幅度调制(AM)和频率调制(FM)包络生成的低频(LF)谱图中的时间信息,采用节奏共振峰分析(RFA)框架,提取前六个节奏共振峰的轨迹特征,并对AM与FM的低频谱图进行二维离散余弦变换(2D-DCT)表征。这些特征输入机器学习模型,用于对比两种语言的节奏模式。该方法实现了无需预先标注语音大单位的实证性节奏结构分析,为东北印度两种低资源语言提供了新的节奏研究范式。
原文摘要 · Abstract (English)
The current work explores long-term speech rhythm variations to classify Mising and Assamese, two low-resourced languages from Assam, Northeast India. We study the temporal information of speech rhythm embedded in low-frequency (LF) spectrograms derived from amplitude (AM) and frequency modulation (FM) envelopes. This quantitative frequency domain analysis of rhythm is supported by the idea of rhythm formant analysis (RFA), originally proposed by Gibbon [1]. We attempt to make the investigation by extracting features derived from trajectories of first six rhythm formants along with two-dimensional discrete cosine transform-based characterizations of the AM and FM LF spectrograms. The derived features are fed as input to a machine learning tool to contrast rhythms of Assamese and Mising. In this way, an improved methodology for empirically investigating rhythm variation structure without prior annotation of the larger unit of the speech signal is illustrated for two low-resourced languages of Northeast India.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。