评测语音大模型性能下降,揭示语音输入导致推理能力衰退。
S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
- 构建语音输入诊断数据集,评估语音大模型的推理与生成能力。
- 通过困惑度差异对比,量化语音输入相比文本输入的性能下降。
- 在百川语音模型上验证,适合研究语音大模型训练与优化者。
端到端语音大语言模型(Speech LLMs)将文本模型能力扩展至直接处理和生成音频标记,但通常会导致其推理与生成性能相较文本输入出现下降,这一现象称为智能退化。为系统评估该差距,我们提出 S2SBench,一个用于量化语音大模型性能退化的基准测试。该基准包含针对句子续写与常识推理的诊断数据集,均在音频输入条件下设计。我们进一步引入基于困惑度差异的成对评估协议,衡量语音输入相对于文本输入的性能退化程度。我们将 S2SBench 应用于百川语音(Baichuan-Audio)的训练过程分析,进一步验证了该基准的有效性。所有数据集与评估代码已公开于 https://github.com/undobug/S2SBench。
原文摘要 · Abstract (English)
End-to-end speech large language models ((LLMs)) extend the capabilities of text-based models to directly process and generate audio tokens. However, this often leads to a decline in reasoning and generation performance compared to text input, a phenomenon referred to as intelligence degradation. To systematically evaluate this gap, we propose S2SBench, a benchmark designed to quantify performance degradation in Speech LLMs. It includes diagnostic datasets targeting sentence continuation and commonsense reasoning under audio input. We further introduce a pairwise evaluation protocol based on perplexity differences between plausible and implausible samples to measure degradation relative to text input. We apply S2SBench to analyze the training process of Baichuan-Audio, which further demonstrates the benchmark's effectiveness. All datasets and evaluation code are available at https://github.com/undobug/S2SBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。