arXiv:2503.06241cs.ROcs.CL2025-03中稿 · presentation at IE…被引 8

提升对话机器人在嘈杂环境中的响应速度,让交互更自然。

A Noise-Robust Turn-Taking System for Real-World Dialogue Robots: A Field Experiment

  • 用Transformer架构设计抗噪语音活动预测模型
  • 实地实验显示响应延迟显著降低,双方反应更快
  • 适合需要真实场景部署的对话机器人研发者

话语交替是人机交互的关键,直接影响对话流畅性和用户参与度。尽管先前研究已在受控环境中探索话语交替模型,但其在真实场景下的鲁棒性仍待深入。本研究提出一种基于Transformer架构的抗噪语音活动投影(VAP)模型,以提升对话机器人的实时话语交替能力。为评估该系统有效性,我们在商场开展了实地实验,对比了VAP系统与传统云端语音识别系统的性能。分析涵盖主观用户评价与客观行为数据。结果表明,所提系统显著降低了响应延迟,使机器人与用户均能更快回应,实现更自然的对话流程。主观评价也显示,更快的响应提升了交互体验。

原文摘要 · Abstract (English)

Turn-taking is a crucial aspect of human-robot interaction, directly influencing conversational fluidity and user engagement. While previous research has explored turn-taking models in controlled environments, their robustness in real-world settings remains underexplored. In this study, we propose a noise-robust voice activity projection (VAP) model, based on a Transformer architecture, to enhance real-time turn-taking in dialogue robots. To evaluate the effectiveness of the proposed system, we conducted a field experiment in a shopping mall, comparing the VAP system with a conventional cloud-based speech recognition system. Our analysis covered both subjective user evaluations and objective behavioral analysis. The results showed that the proposed system significantly reduced response latency, leading to a more natural conversation where both the robot and users responded faster. The subjective evaluations suggested that faster responses contribute to a better interaction experience.

对话系统语音识别抗噪处理真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。