构建中文语音交互评估基准,检验模型在真实听觉情境下的自然应答能力。
TELEVAL: A Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios
- 设计无指令、音频驱动的评测场景,模拟真实语音交互环境。
- 发现模型在声学变化下表现下降,存在描述音频而非互动的共性缺陷。
- 提出'画面陷阱'现象,适合研究语音交互与听觉上下文对齐的学者使用。
语音语言模型(SLMs)应支持超越任务完成的自然口语交互。然而,现有基准主要评估结构化场景中的语义正确性,对基于声学上下文的交互行为评估有限。为此,我们引入TELEVAL,一个大规模的中文语音交互评测基准,适用于无指令、音频条件下的评测场景。TELEVAL评估两个互补维度:(1) 可靠内容实现,衡量模型在多样化声学与语言条件下语义准确性;(2) 交互适当性,评估模型是否通过隐式依赖听觉线索生成自然且恰当的回应。实验表明,尽管模型在语义任务上表现良好,但在声学变异和交互场景中性能显著下降。观察到从感知不稳到交互错误的持续退化,并识别出一种重复出现的失败模式,称为“画面陷阱”——模型倾向于描述感知到的音频信号,而非生成适当的交互响应。这些结果表明,当前SLMs仍未能充分满足自然口语交互的需求。TELEVAL为评估和分析SLMs的交互行为提供了针对性框架。
原文摘要 · Abstract (English)
Spoken Language Models (SLMs) are expected to support natural spoken interaction beyond task completion. However, existing SLM benchmarks primarily evaluate semantic correctness in structured settings and provide limited assessment of interactional behavior grounded in acoustic context. To address this gap, we introduce TELEVAL, a large-scale SLM benchmark for Chinese spoken interaction in instruction-free, audio-conditioned settings. TELEVAL evaluates two complementary aspects: (1) Reliable Content Fulfillment, which measures semantic accuracy of SLMs under diverse acoustic and linguistic conditions, and (2) Interactional Appropriateness, which assesses whether models produce natural and appropriate responses by implicitly grounding behavior in auditory cues. Experiments show that while models perform competitively on semantic tasks, their performance degrades under acoustic variability and in interactional settings. We observe consistent degradation from perceptual instability to interactional errors, and further identify a recurring failure pattern, termed the "Caption Trap", where models tend to describe perceived audio signals rather than produce appropriate interactive responses. These results indicate that current SLMs remain insufficiently aligned with the requirements of natural spoken interaction. TELEVAL provides a targeted framework for evaluating and analyzing interactional behavior in SLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。