arXiv:2505.19103cs.CLcs.SD2025-05中稿 · Interspeech2025被引 7

让语音转录更精准,自动识别句子中强调的词语。

WHISTRESS: Enriching Transcriptions with Sentence Stress Detection

  • 无需对齐数据,直接从语音中检测句子重音位置。
  • 在合成数据上训练,仍能在多个真实场景零样本泛化。
  • 适合需要语义准确转录的语音系统开发者。

口语不仅通过词汇传递意义,还依赖语调、情感和强调。句子重音是句中特定词语的强调,对表达说话人意图至关重要,已在语言学中广泛研究。本文提出 WHISTRESS,一种无需对齐的句子重音检测方法,用于增强转录系统。为此,我们构建了 TINYSTRESS-15K——一个可扩展的合成训练数据集,通过完全自动化流程生成。在该数据集上训练 WHISTRESS 后,其性能优于现有方法,且训练与推理均无需额外先验输入。值得注意的是,尽管仅在合成数据上训练,WHISTRESS 在多个不同基准上展现出强大的零样本泛化能力。

原文摘要 · Abstract (English)

Spoken language conveys meaning not only through words but also through intonation, emotion, and emphasis. Sentence stress, the emphasis placed on specific words within a sentence, is crucial for conveying speaker intent and has been extensively studied in linguistics. In this work, we introduce WHISTRESS, an alignment-free approach for enhancing transcription systems with sentence stress detection. To support this task, we propose TINYSTRESS-15K, a scalable, synthetic training data for the task of sentence stress detection which resulted from a fully automated dataset creation process. We train WHISTRESS on TINYSTRESS-15K and evaluate it against several competitive baselines. Our results show that WHISTRESS outperforms existing methods while requiring no additional input priors during training or inference. Notably, despite being trained on synthetic data, WHISTRESS demonstrates strong zero-shot generalization across diverse benchmarks. Project page: https://pages.cs.huji.ac.il/adiyoss-lab/whistress.

语音转录重音检测合成数据零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。