arXiv:2601.13835cs.CL2026-01中稿 · ICASSP 2026

探究自监督语音表征中韵律与词汇线索对对话轮换的作用

The Role of Prosodic and Lexical Cues in Turn-Taking with Self-Supervised Speech Representations

  • 用声码器分离控制语音的韵律和词汇信息,精准测试模型依赖
  • 仅凭韵律即可实现接近原始语音的轮换预测准确率
  • 模型能自动切换依赖线索,适合隐私敏感或资源受限场景

流畅的对话轮换仍是人机交互中的关键挑战。自监督语音表征(S3Rs)推动了诸多进展,但基于S3R的轮换模型是否依赖韵律、词汇或两者兼有仍不明确。本文提出一种基于声码器的方法,更清晰地控制语音的韵律与词汇线索,从而探测一个基于S3R的语音活动投影模型。结果发现,在韵律匹配但语义不可懂的噪声语音上,模型预测准确率与在清晰语音上相近。这表明韵律与词汇线索均支持轮换,且任一可独立使用。因此未来模型或仅需韵律,兼具隐私保护与性能优势。当其中任一信息被干扰时,模型能自动利用另一线索,无需额外训练,说明二者在S3Rs中编码程度低、依赖弱。结果在基于CPC与wav2vec2.0的S3Rs中均一致。所有代码均已公开,支持后续研究。

原文摘要 · Abstract (English)

Fluid turn-taking remains a key challenge in human-robot interaction. Self-supervised speech representations (S3Rs) have driven many advances, but it remains unclear whether S3R-based turn-taking models rely on prosodic cues, lexical cues or both. We introduce a vocoder-based approach to control prosody and lexical cues in speech more cleanly than prior work. This allows us to probe the voice-activity projection model, an S3R-based turn-taking model. We find that prediction on prosody-matched, unintelligible noise is similar to accuracy on clean speech. This reveals both prosodic and lexical cues support turn-taking, but either can be used in isolation. Hence, future models may only require prosody, providing privacy and potential performance benefits. When either prosodic or lexical information is disrupted, the model exploits the other without further training, indicating they are encoded in S3Rs with limited interdependence. Results are consistent in CPC-based and wav2vec2.0 S3Rs. We discuss our findings and highlight a number of directions for future work. All code is available to support future research.

语音交互自监督学习对话系统韵律分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。