现有语音大模型识别说话人能力弱,依赖文字而非声音作答。
Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue
- 通过对比文本与音频信息,发现模型更依赖对话转录文本。
- 在身份关键问题上准确率显著低于上下文类问题。
- 建议用需识别人物身份的任务评估语音大模型真实能力。
近年来,语音大语言模型(SpeechLLMs)快速发展,接近人类听觉理解与推理能力。在高考英语听力等基准测试中表现优异,看似需同时理解话语内容与说话人特征。但经分析发现,多数题目仅凭对话转录文本即可推断答案,无需说话人分割与识别。对Qwen-Audio和WavLLM在高考及自建的'What Do You Like?'数据集上的评估显示,模型在上下文类问题上准确率远高于需说话人身份的关键问题。结果表明,当前SpeechLLMs在解决语音问答任务时,对音频中的说话人信息感知有限,行为类似仅基于文本的LLM推理。我们提出,应采用聚焦身份关键问题的任务,以更真实评估SpeechLLMs在语音问答中的能力。
原文摘要 · Abstract (English)
In recent years, we have observed a rapid advancement in speech language models (SpeechLLMs), catching up with humans' listening and reasoning abilities. SpeechLLMs have demonstrated impressive spoken dialog question-answering (SQA) performance in benchmarks like Gaokao, the English listening test of the college entrance exam in China, which seemingly requires understanding both the spoken content and voice characteristics of speakers in a conversation. However, after carefully examining Gaokao's questions, we find the correct answers to many questions can be inferred from the conversation transcript alone, i.e.\ without speaker segmentation and identification. Our evaluation of state-of-the-art models Qwen-Audio and WavLLM on both Gaokao and our proposed "What Do You Like?" dataset shows a significantly higher accuracy in these context-based questions than in identity-critical questions, which can only be answered reliably with correct speaker identification. The results and analysis suggest that when solving SQA, the current SpeechLLMs exhibit limited speaker awareness from the audio and behave similarly to an LLM reasoning from the conversation transcription without sound. We propose that tasks focused on identity-critical questions could offer a more accurate evaluation framework of SpeechLLMs in SQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。