arXiv:2602.15519eess.AScs.SD2026-02中稿 · Interspeech 2026

用唤醒词自动做语音录入,让人机对话更自然。

Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios

  • 用交互时的唤醒词作语音录入参考,无需提前录音
  • 在真实嘈杂环境下,语音识别准确率仍有提升空间
  • 用大模型生成语音增强录入,显著改善听感

目标语音提取(TSE)通常依赖于预先录制的高质量语音样本,这会打断用户体验并限制在自发交互场景中的可行性。本文提出一种名为 Enroll-on-Wakeup(EoW)的新框架,将用户与机器交互过程中自然捕获的唤醒词片段自动作为注册参考。该方法消除了对预采集语音的需求,实现了无缝体验。我们首次系统性地研究了 EoW-TSE,评估了先进判别式与生成式模型在真实多样声学条件下的表现。由于唤醒词片段短且噪声大,我们探索了使用基于大语言模型的文本转语音(LLM-based TTS)进行语音增强。结果表明,尽管当前 TSE 模型在 EoW-TSE 场景下性能下降,但借助 TTS 辅助可显著提升听觉体验,语音识别准确率方面仍存在差距。

原文摘要 · Abstract (English)

Target speech extraction (TSE) typically relies on pre-recorded high-quality enrollment speech, which disrupts user experience and limits feasibility in spontaneous interaction. In this paper, we propose Enroll-on-Wakeup (EoW), a novel framework where the wake-word segment, captured naturally during human-machine interaction, is automatically utilized as the enrollment reference. This eliminates the need for pre-collected speech to enable a seamless experience. We perform the first systematic study of EoW-TSE, evaluating advanced discriminative and generative models under real diverse acoustic conditions. Given the short and noisy nature of wake-word segments, we investigate enrollment augmentation using LLM-based TTS. Results show that while current TSE models face performance degradation in EoW-TSE, TTS-based assistance significantly enhances the listening experience, though gaps remain in speech recognition accuracy.

语音提取人机交互唤醒词TTS增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。