arXiv:2608.27783eess.AScs.CL2026-08

提出语音模型准入评估挑战,发现回答前的错误检测盲区。

SURE-Challenge: Evaluating Speech Evidence Before Speech-LLM Generation

  • 设计新基准测试语音模型准入决策,区分有效与无效输入
  • 固定规则可识别196/204个无效输入,准确率不降
  • 适合关注语音模型鲁棒性与前置过滤机制的研究者

语音大模型通常在生成回答后才进行评估,但系统需先决定是否将音频波形送入模型。为此,我们提出语音支持性拒绝评估挑战(SURE-Challenge)用于该准入环节。基准测试在独立数据源划分下,配对了来自LibriSpeech的转录文本与首词问答,以及无支持沉默、有色噪声、合成音调和源模糊人声干扰。前端实验使用Qwen2-Audio;选定的能量+Whisper评分规则在六种语音/音频大模型前重演。在经泄露筛选的474样本测试集上,原始Qwen2-Audio仅拒掉15/204个无效输入,而固定规则拒掉196/204,且不影响有效输入的准确率。外部评估进一步验证:随着Whisper评分阈值收紧,Common Voice保留率下降;无速度人声干扰在54段再生种子中导致18至24段被拒。结果揭示了仅靠回答评分会忽略的生成前错误模式。

原文摘要 · Abstract (English)

Speech LLMs are usually graded after they answer, although an operating system first has to decide whether to send a waveform to the model. We define the Speech-Unsupported Rejection Evaluation Challenge (SURE-Challenge) for this admission step. The benchmark pairs LibriSpeech-derived transcription and first-word question answering with unsupported silence, colored noise, synthetic tones, and source-ambiguous babble under disjoint source splits. Front-end ablations use Qwen2-Audio; the selected energy-plus-Whisper-score rule is then replayed before six speech/audio LLMs. On the leakage-screened 474-example SURE-Extended test set, raw Qwen2-Audio rejects 15/204 unsupported inputs, whereas the fixed rule rejects 196/204 and leaves supported accuracy unchanged. External evaluations qualify this result: Common Voice retention drops as the Whisper-score threshold is tightened, and no-speed babble gives 18 to 24 rejected clips out of 54 across regenerated seeds. The result identifies a pre-generation error mode missed by answer-only scoring.

语音模型评估基准前置过滤

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。