arXiv:2607.05365cs.CLcs.AI2026-07

构建对话自然度评估基准,揭示语音翻译模型在交互细节上的不足。

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

论文配图:SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models
图 1 · 摘自论文原文
  • 基于可控对话数据,多维度评估语音生成模型表现
  • 现有模型虽保真度高、识别准,但交互延迟与情感适配仍差
  • 适合语音交互系统研发者与评估人员参考

流式语音到语音语言模型旨在直接以合成语音回应口语提问。然而,标准的语音与文本评估基准无法捕捉这些系统在对话中是否自然,而自然性由时机、轮流发言、语调、人际立场、语言与方言一致性以及关系感知的恰当性共同决定。我们提出SPEARBench,一个针对语音到语音语言模型在问答交互中自然度的评估基准。该基准从Seamless Interaction语料库构建受控对话提示,对多个模型进行推理,并采用多维度评估协议,涵盖响应延迟、打断、语音质量、语音识别鲁棒性、语言与方言一致性、情感自然度、人际立场及可解释的分布基线。基准包含原始人类回答作为参考条件,并报告了多个当代模型的结果。结果显示,当前模型虽能达到高信号级质量与低语音识别错误率,但在响应延迟、重叠、方言保持、情感适应和人际立场动态方面仍与人类对话行为存在差异。

原文摘要 · Abstract (English)

Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, standard speech and text benchmarks do not capture whether these systems behave naturally in conversations, where timing, turn-taking, prosody, interpersonal stance, language and dialect consistency, and relationship-aware appropriateness jointly shape perceived quality. We introduce SPEARBench, a benchmark for evaluating naturalness in speech-to-speech language models from question-answer interactions. SPEARBench constructs controlled dialogue prompts from the Seamless Interaction corpus, runs inference across multiple models, and evaluates generated answers using a multidimensional protocol that covers response latency, interruptions, speech quality, ASR robustness, language and dialect consistency, emotional naturalness, interpersonal stance, and explainable distributional baselines. The benchmark includes original human answers as a reference condition and reports results for several contemporary models. Results show that current models can achieve high signal-level quality and low ASR error while still differing from human conversational behavior in latency, overlap, dialect preservation, emotional adaptation, and interpersonal stance dynamics.

语音生成对话评估自然度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。