arXiv:2505.22236cs.CL2025-05中稿 · CoNLL 2025被引 1

研究语音合成对语法边界的敏感度,发现系统在复杂句中依赖标点而非语义。

A Linguistically Motivated Analysis of Intonational Phrasing in Text-to-Speech Systems: Revealing Gaps in Syntactic Sensitivity

  • 用心理语言学方法分析语音合成中的语调切分机制
  • 复杂句中系统准确率下降,需依赖逗号等表面标记
  • 微调后模型更依赖深层语义,生成更自然的语调模式

我们采用心理语言学启发的方法,分析文本到语音(TTS)系统在语调切分上的句法敏感性。语调边界通常可由句子内的句法边界预测。研究发现,当句法边界模糊时(如花园路径句或附着歧义句),TTS系统难以准确生成语调边界,需依赖逗号等表面提示;而在句法结构简单的句子中,系统能有效利用句法线索,超越表层标记。随后,我们对无逗号标注句法边界的句子进行微调,促使模型关注更细微的语言线索。结果表明,该策略使语调模式更具区分性,更好地反映深层句法结构。

原文摘要 · Abstract (English)

We analyze the syntactic sensitivity of Text-to-Speech (TTS) systems using methods inspired by psycholinguistic research. Specifically, we focus on the generation of intonational phrase boundaries, which can often be predicted by identifying syntactic boundaries within a sentence. We find that TTS systems struggle to accurately generate intonational phrase boundaries in sentences where syntactic boundaries are ambiguous (e.g., garden path sentences or sentences with attachment ambiguity). In these cases, systems need superficial cues such as commas to place boundaries at the correct positions. In contrast, for sentences with simpler syntactic structures, we find that systems do incorporate syntactic cues beyond surface markers. Finally, we finetune models on sentences without commas at the syntactic boundary positions, encouraging them to focus on more subtle linguistic cues. Our findings indicate that this leads to more distinct intonation patterns that better reflect the underlying structure.

语音合成语调建模句法敏感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。