arXiv:2604.25309eess.ASeess.SP2026-04

用声学节奏与频谱特征分析印度两种濒危语言的差异

Cross-Linguistic Rhythmic and Spectral Feature-Based Analysis of Nyishi and Adi: Two Under-Resourced Languages of Arunachal Pradesh

  • 通过低频调制谱分析语音节奏,提取三类韵律特征
  • 两类语言在节奏和频谱上呈现层级分化,分类准确率达93.96%
  • 适合研究语言多样性、语音识别及濒危语言保护的学者

濒危语言在定量节奏研究中仍被忽视,尤其缺乏对密切相关的语言群体内部声学差异的系统分析。本研究以印度东北部阿鲁纳恰尔邦的尼西语和阿迪语为对象,采用基于幅度调制低频谱分析(即节奏共振峰分析,RFA)的频域框架,探究蒂尼语支内部的语言声学差异。从低频调制谱中提取三个节奏共振峰特征:主导峰数量(NDP)、主导峰平均频率(MFDP)和主导频率方差(VFDP)。同时提取离散余弦变换(DCT)系数和梅尔频率倒谱系数(MFCC),刻画语音信号的谱调制结构与整体频谱组织。统计建模揭示出层级分化模式:节奏特征表现出一致但中等程度的分离,尼西语具有更高的主导调制频率和更大的频率离散度;分类实验进一步验证该层次关系,仅使用节奏特征即可达到约84-85%的分类准确率,融合MFCC后提升至支持向量机(SVM)90.9%和多层感知机(MLP)93.96%。结果表明,节奏与频谱特征编码互补的語言变异,低频调制捕捉宏观时间结构,而频谱特征反映更精细的音系分化。

原文摘要 · Abstract (English)

Under-resourced languages remain underrepresented in quantitative rhythm research,particularly in systematic intra-branch analysis of acoustic differentiation within closely related linguistic groups.This study investigates acoustic differentiation within the Tani language subgroup by examining speech rhythm in Nyishi and Adi,two under-resourced Tani languages spoken in Arunachal Pradesh,North-East India,using a frequency domain framework based on amplitude modulation(AM) low-frequency(LF) spectrum analysis,commonly referred to as rhythm formant analysis(RFA).The analysis is designed to identify whether intra-branch differentiation follows a hierarchical pattern across rhythmic and spectral domains.From the LF modulation spectrum,three rhythm formant features were derived:Number of Dominant peaks(NDP),Mean Frequency of Dominant Peaks(MFDP),and Variance of Dominant Frequencies(VFDP).In addition,Discrete Cosine Transform (DCT)coefficients and Mel Frequency Cepstral Coefficient(MFCC) were extracted to characterise the spectral modulation structure and broad spectral organisation of the speech signal.Statistical modelling reveals a hierarchical pattern of differentiation,where rhythmic features show consistent but moderate separation,with Nyishi exhibiting higher dominant modulation frequencies as well as greater dispersion than Adi.Classification experiments further support this hierarchy,with rhythm-only features achieved approximately 84-85% classification accuracy.Fusion using MFCC representations improved performance to 90.9% classification accuracy using support vector machine (SVM) and 93.96% using multilayer perceptron (MLP).These findings demonstrate that rhythmic and spectral features encode complementary levels of linguistic variations,with low frequency modulation capturing constrained macro temporal structure and spectral features reflecting finer phonological differentiation.

语音分析语言多样性濒危语言节奏特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。