arXiv:2502.05389cs.CL2025-02NAACL被引 7

探究语音语调在口语问答中的作用,发现其虽有用但常被文字信息掩盖。

The Role of Prosody in Spoken Question Answering

  • 分离语音语调与词汇信息,在真实语音数据上测试模型表现。
  • 仅用语调信息训练的模型也能达到合理准确率,说明语调含有效信息。
  • 当词汇信息可用时,模型几乎只依赖词汇,需改进融合方法提升语调作用。

目前的口语理解研究大多带有强烈的文本视角:多数数据集源自文本合成语音,模型也依赖语音转写文本。这导致语音中除发音外的附加信息——语调——被忽视,而语调难以从文本中恢复。本文在包含自然语音的SLUE-SQA-5数据集上,分离语调与词汇信息,研究其在口语问答中的作用。结果表明,仅使用语调信息训练的模型仍能取得较好性能,说明语调蕴含有价值线索;但当词汇信息存在时,模型倾向于完全依赖词汇。研究提示:尽管语调可提供补充信息,但需更有效的融合机制,才能使其在与词汇特征并行时发挥更大作用。

原文摘要 · Abstract (English)

Spoken language understanding research to date has generally carried a heavy text perspective. Most datasets are derived from text, which is then subsequently synthesized into speech, and most models typically rely on automatic transcriptions of speech. This is to the detriment of prosody--additional information carried by the speech signal beyond the phonetics of the words themselves and difficult to recover from text alone. In this work, we investigate the role of prosody in Spoken Question Answering. By isolating prosodic and lexical information on the SLUE-SQA-5 dataset, which consists of natural speech, we demonstrate that models trained on prosodic information alone can perform reasonably well by utilizing prosodic cues. However, we find that when lexical information is available, models tend to predominantly rely on it. Our findings suggest that while prosodic cues provide valuable supplementary information, more effective integration methods are required to ensure prosody contributes more significantly alongside lexical features.

语音理解语调分析问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。