arXiv:2607.02002cs.CL2026-07被引 1

用上下文嵌入预测普通话单音节词发音时长和音高,准确度显著高于随机水平。

Using embeddings to predict spoken word duration and pitch in Mandarin monosyllabic words

论文配图:Using embeddings to predict spoken word duration and pitch in Mandarin monosyllabic words
图 1 · 摘自论文原文
  • 利用上下文嵌入预测单音节词发音时长,效果优于随机猜测。
  • 预测时长精度足够还原音高轮廓到毫秒级时间尺度。
  • 适用于语音合成与自然语言处理中的韵律建模研究。

普通话对话中单音节词的时间归一化音高轮廓部分可由其上下文嵌入(CEs)预测。本研究分析了从自发性普通话语料库中提取的7470个单音节CV词的发音时长,发现上下文嵌入在词类层面及个体词层面均具有显著预测能力,经类型与实例级置换基线验证。预测时长具备足够精度,可将[0,1]归一化时间尺度下的音高轮廓转换为毫秒级真实时间轮廓,生成结果逼近实际轮廓,且优于置换基线。

原文摘要 · Abstract (English)

Time-normalized f0 contours of Mandarin words in conversational speech have been shown to be predictable in part from their contextualized embeddings (CEs). The present study investigates whether CEs also predict spoken word duration for 7470 tokens of Mandarin monosyllabic CV words extracted from a Mandarin corpus of spontaneous speech. We show that CEs indeed are predictive for duration, above chance level, not only at the type level, but also at the level of individual tokens, as indicated by the results obtained with the type-wise and token-wise permutation baselines. We also show that the predicted durations are sufficiently precise to back-transform predicted f0 contours in [0,1] normalized time to contours on the ms time scale. The resulting predicted contours approximate empirical contours and also outperform a permutation baseline.

语音合成嵌入模型韵律预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。