arXiv:2509.26388eess.AScs.AI2025-09中稿 · ICASSP 2026被引 16

提出新基准,评估语音模型的实时交互能力

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models

  • 设计游戏化任务评估语音模型的时间感知能力
  • 多数模型在节奏和同步任务中表现显著下降
  • 适合研究实时语音交互与多轮对话的团队

对话式语音语言模型(SLMs)正成为实时语音交互的新兴范式。然而,其在时间动态性方面的能力——包括时间管理、语速控制和同时说话的处理——仍是影响对话流畅性的关键未评估挑战。为填补这一空白,我们提出了Game-Time基准,一个系统评估这些时间能力的框架。受人类通过语言活动习得语言的启发,Game-Time包含基础指令跟随任务和带有时间约束的高级任务,如节奏遵循和同步响应。对多种SLM架构的评估显示显著性能差异:虽然顶尖模型能较好完成基础任务,但许多现有系统仍难以应对基本指令;更关键的是,几乎所有模型在时间约束下表现大幅退化,暴露出对时间意识和全双工交互的持续弱点。Game-Time基准为未来研究提供了方向。演示与数据集见项目网站 https://ga642381.github.io/Game-Time。

原文摘要 · Abstract (English)

Conversational Spoken Language Models (SLMs) are emerging as a promising paradigm for real-time speech interaction. However, their capacity of temporal dynamics, including the ability to manage timing, tempo and simultaneous speaking, remains a critical and unevaluated challenge for conversational fluency. To address this gap, we introduce the Game-Time Benchmark, a framework to systematically assess these temporal capabilities. Inspired by how humans learn a language through language activities, Game-Time consists of basic instruction-following tasks and advanced tasks with temporal constraints, such as tempo adherence and synchronized responses. Our evaluation of diverse SLM architectures reveals a clear performance disparity: while state-of-the-art models handle basic tasks well, many contemporary systems still struggle with fundamental instruction-following. More critically, nearly all models degrade substantially under temporal constraints, exposing persistent weaknesses in time awareness and full-duplex interaction. The Game-Time Benchmark provides a foundation for guiding future research toward more temporally-aware conversational AI. Demos and datasets are available on our project website https://ga642381.github.io/Game-Time.

语音模型实时交互时间动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。