arXiv:2604.10580cs.CLcs.SD2026-04

测试语音合成系统能否根据语境正确强调关键词

Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark

  • 设计对比语境对,让相同句子在不同上下文中需强调不同词
  • 发现纯文本模型能准确推断应强调的词,但语音合成系统常失败
  • 适合研究上下文感知语音合成的学者和开发者

口语意义不仅取决于说了什么,还取决于哪个词被强调。同一句话在不同强调位置下可传达纠正、对比或澄清等含义。尽管现代文本转语音(TTS)系统能生成有表现力的语音,但尚不清楚它们是否能仅凭语境推断出合适的重音。为此,我们提出上下文感知重音语音合成(CAST)基准,用于评估TTS系统在语境条件下对词级重音的处理能力。该基准采用对比性语境对:相同句子搭配不同语境,要求强调不同的词。我们评估了当前最先进的系统,发现一致差距:仅依赖文本的语言模型能可靠地从语境中恢复预期重音,而TTS系统在语音实现上却频繁失败。我们公开了该基准、评估框架、构建流程及一个合成语料库,以支持未来上下文感知语音合成的研究。

原文摘要 · Abstract (English)

Spoken meaning often depends not only on what is said, but also on which word is emphasized. The same sentence can convey correction, contrast, or clarification depending on where emphasis falls. Although modern text-to-speech (TTS) systems generate expressive speech, it remains unclear whether they infer contextually appropriate stress from discourse alone. To address this gap, we present Context-Aware Stress TTS (CAST), a benchmark for evaluating context-conditioned word-level stress in TTS. Items are defined as contrastive context pairs: identical sentences paired with distinct contexts requiring different stressed words. We evaluate state-of-the-art systems and find a consistent gap: text-only language models reliably recover the intended stress from context, yet TTS systems frequently fail to realize it in speech. We release the benchmark, evaluation framework, construction pipeline and a synthetic corpus to support future work on context-aware speech synthesis.

语音合成上下文感知重音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。