arXiv:2603.16783cs.CL2026-03中稿 · EMNLP

构建可模拟真实语音交互的用户模型,提升语音对话系统评估效果。

SpokenUS: A Spoken User Simulator for Task-Oriented Dialogue

  • 基于5万+对话数据设计语音用户模拟器,支持跨轮槽位、插话等四类真实语音行为。
  • 在1034小时语音上测试,目标完成率媲美更大模型,人类评分显著更高。
  • 适合用于训练和评估更鲁棒的语音对话系统,尤其关注自然交互表现。

鲁棒的语音代理需要接触多样化的实际语音交互,但真实语音数据获取成本高昂。现有任务导向对话(TOD)数据集规模小、领域覆盖有限,且缺乏系统化扩充方法。为此,我们提出SpokenTOD,一个包含52,390条对话和1,034小时语音的大规模语音任务导向对话数据集,涵盖跨轮槽位、插话、不流畅表达和情感语调四种真实语音行为,覆盖多样说话人与领域。在此基础上,我们构建了SpokenUS,一种基于TOD的语音用户模拟器,通过专用的说话时机决策模块控制发言节奏。SpokenUS在目标覆盖率上达到远超自身规模的基线水平,且在人类感知评分(MOS)上显著优于所有基线,在对话中逐步披露槽位值,模仿人类行为而非一次性暴露。进一步分析表明,SpokenUS所模拟的语音行为对语音代理构成真实挑战,使其成为评估更鲁棒语音对话系统的实用工具。代码已开源。

原文摘要 · Abstract (English)

Robust voice agents require exposure to the full diversity of how people interact through speech. However, obtaining enough spoken interactions is prohibitively expensive. Building spoken user simulators that address this requires large-scale spoken task-oriented dialogue (TOD) data encompassing spoken user behaviors, yet existing datasets are limited in scale and domain coverage, with no systematic pipeline for augmenting them. To address this, we introduce SpokenTOD, a spoken TOD dataset of 52,390 dialogues and 1,034 hours of speech augmented with four spoken user behaviors---cross-turn slots, barge-in, disfluency, and emotional prosody---across diverse speakers and domains. Building on SpokenTOD, we present SpokenUS, a spoken user simulator grounded in TOD that decides when to speak through a dedicated turn-taking head. SpokenUS achieves comparable goal coverage to much larger models while substantially outperforming all baselines in human MOS, disclosing slot values gradually across the dialogue as humans do rather than front-loading them. Further analysis confirms that SpokenUS's spoken behaviors pose meaningful challenges to voice agents, making it a practical tool for evaluating more robust spoken dialogue systems. Our code is available at https://github.com/holi-lab/SpokenUS.

语音对话用户模拟任务导向数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。