arXiv:2503.04721cs.CLeess.AS2025-03中稿 · ASRU 2025被引 87

评测对话模型实时交互能力,推动更自然的语音对话发展。

Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

  • 构建全双工对话评估基准,聚焦停顿、回应、抢话等交互行为。
  • 采用自动指标实现快速、可复现的评估,提升评测效率与公平性。
  • 适合研究语音对话系统、人机交互的开发者和研究人员使用。

语音对话建模面临文本语言模型之外的挑战,需支持实时交互、换言和反馈语用。当前多数语音对话模型(SDMs)采用半双工模式,一次处理一个对话轮次;而新兴的全双工模型可同时听与说,使对话更自然。然而现有评估仍局限于轮次级指标或粗粒度语料分析。为此,我们提出 Full-Duplex-Bench,一个系统化评估关键交互行为的基准,涵盖停顿处理、反馈语用、换言策略与打断管理。该框架采用自动指标,确保评估的一致性与可复现性,并提供公平、高效的评测环境。通过公开基准与代码,旨在推动语音对话建模发展,促进更自然、更具互动性的对话模型研发。

原文摘要 · Abstract (English)

Spoken dialogue modeling poses challenges beyond text-based language modeling, requiring real-time interaction, turn-taking, and backchanneling. While most Spoken Dialogue Models (SDMs) operate in half-duplex mode-processing one turn at a time - emerging full-duplex SDMs can listen and speak simultaneously, enabling more natural conversations. However, current evaluations remain limited, focusing mainly on turn-based metrics or coarse corpus-level analyses. To address this, we introduce Full-Duplex-Bench, a benchmark that systematically evaluates key interactive behaviors: pause handling, backchanneling, turn-taking, and interruption management. Our framework uses automatic metrics for consistent, reproducible assessment and provides a fair, fast evaluation setup. By releasing our benchmark and code, we aim to advance spoken dialogue modeling and foster the development of more natural and engaging SDMs.

语音对话全双工评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。