用语音识别模型量化藏语声调演化阶段,发现声调功能随语言演变渐变。
The Tonogenesis Continuum in Tibetan: A Computational Investigation
- 用语音识别模型测试不同藏语方言对声调变化的敏感度
- 发现阿莫方言最不依赖声调,乌斯藏方言最依赖,喀木方言居中
- 揭示声调演化是渐进过程,适合研究语言变迁与语音系统转型
声调演化——即音段对立向词汇声调转变的历史过程——传统上通过比较重构和声学语音学研究。本文提出一种计算方法,通过测量声调操作对自动语音识别(ASR)性能的影响,量化不同阶段声调的功能作用。基于一组密切相关的藏语方言,分析其对声调平坦化的敏感性,发现声调演化存在连续谱:无声调的阿莫方言对声调移除容忍度最高,完全声调化的乌斯藏方言则出现严重性能下降,而中间的喀木方言表现介于两者之间。这种梯度效应表明,ASR模型能隐式学习到声调功能负荷随语言从辅音主导转向声调主导的动态变化。研究结果说明计算方法可捕捉音变的细微阶段,提示仅依赖最小对立对的传统功能负荷指标,可能高估了过渡系统中声调的依赖程度,因为音段与超音段线索在发音上仍紧密交织。
原文摘要 · Abstract (English)
Tonogenesis-the historical process by which segmental contrasts evolve into lexical tone-has traditionally been studied through comparative reconstruction and acoustic phonetics. We introduce a computational approach that quantifies the functional role of pitch at different stages of this sound change by measuring how pitch manipulation affects automatic speech recognition (ASR) performance. Through analysis on the sensitivity to pitch-flattening from a set of closely related Tibetan languages, we find evidence of a tonogenesis continuum: atonal Amdo dialects tolerate pitch removal the most, while fully tonal U-Tsang varieties show severe degradation, and intermediate Kham dialects fall measurably between these extremes. These gradient effects demonstrate how ASR models implicitly learn the shifting functional load of pitch as languages transition from consonant-based to tone-based lexical contrasts. Our findings show that computational methods can capture fine-grained stages of sound change and suggest that traditional functional load metrics, based solely on minimal pairs, may overestimate pitch dependence in transitional systems where segmental and suprasegmental cues remain phonetically intertwined.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。