arXiv:2503.23163cs.CL2025-03综述被引 10

语境中的词义显著影响台湾国语发音,研究发现其比声调模式更关键。

The realization of tones in spontaneous spoken Taiwan Mandarin: a corpus-based survey and theory-driven computational modeling

  • 用广义加性混合模型分析20种声调组合的语音数据,识别关键影响因素
  • 词义和上下文嵌入能解释发音中超过声调模式的影响,效果更显著
  • 适合对语音与语义关系、计算语言学感兴趣的学者参考

大量文献表明语义可共同决定精细语音特征,但语音实现与语义之间的复杂互动仍研究不足,尤其在声调实现方面。本研究基于台湾国语自发口语语料库,考察了包含全部20种双音节词声调组合的声调实现情况。采用广义加性混合模型(GAMs)将基频轮廓建模为性别、声调语境、声调模式、语速、词位置、二元语法概率、说话人及词汇等预测变量的函数。结果显示,词汇和语义是基频轮廓的关键预测因子,其效应量超过声调模式。对数据集中每个词项,利用GPT-2大语言模型在其语境中生成上下文嵌入表示。结果表明,这些上下文嵌入可相当程度上预测词项的基频轮廓,近似于使用情境下的词义特异性。研究表明,语义与语音实现的纠缠远超传统语言理论预期。

原文摘要 · Abstract (English)

A growing body of literature has demonstrated that semantics can co-determine fine phonetic detail. However, the complex interplay between phonetic realization and semantics remains understudied, particularly in pitch realization. The current study investigates the tonal realization of Mandarin disyllabic words with all 20 possible combinations of two tones, as found in a corpus of Taiwan Mandarin spontaneous speech. We made use of Generalized Additive Mixed Models (GAMs) to model f0 contours as a function of a series of predictors, including gender, tonal context, tone pattern, speech rate, word position, bigram probability, speaker and word. In the GAM analysis, word and sense emerged as crucial predictors of f0 contours, with effect sizes that exceed those of tone pattern. For each word token in our dataset, we then obtained a contextualized embedding by applying the GPT-2 large language model to the context of that token in the corpus. We show that the pitch contours of word tokens can be predicted to a considerable extent from these contextualized embeddings, which approximate token-specific meanings in contexts of use. The results of our corpus study show that meaning in context and phonetic realization are far more entangled than standard linguistic theory predicts.

语音学语义声调上下文嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。