测试语音大模型对口语冗余的处理能力,发现训练目标决定其鲁棒性。
Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones
- 用口语中的不连贯内容作为测试,评估模型修复结构的能力。
- 不同模型表现稳定分群,推理型模型过度删除流畅内容。
- 微调虽提升效果却降低泛化能力,训练目标影响鲁棒性。
语音大模型依赖语言模型作为核心,但其对自发对话输入的行为仍不清楚。口语中普遍存在不连贯现象——插话、修改和插入语,这些在预训练所用书面语料中极为罕见。由于黄金标准的不连贯去除是仅删除的任务,可作为控制实验来判断模型是忠实修复结构还是存在偏见的重新解释。我们使用DRES评估框架,在多种架构与规模的专有及开源语言模型上进行测试。结果表明,模型性能形成稳定的精确率-召回率区域,反映不同的编辑策略。值得注意的是,推理类模型系统性地过度删除流畅内容,暴露出对语义抽象而非结构保真的偏好。尽管微调能达到当前最佳表现,但损害了泛化能力。研究揭示,对语音的鲁棒性由特定训练目标塑造。
原文摘要 · Abstract (English)
LLMs serve as the backbone in SpeechLLMs, yet their behavior on spontaneous conversational input remains poorly understood. Conversational speech contains pervasive disfluencies -- interjections, edits, and parentheticals -- that are rare in the written corpora used for pre-training. Because gold disfluency removal is a deletion-only task, it serves as a controlled probe to determine whether a model performs faithful structural repair or biased reinterpretation. Using the DRES evaluation framework, we evaluate proprietary and open-source LLMs across architectures and scales. We show that model performance clusters into stable precision-recall regimes reflecting distinct editing policies. Notably, reasoning models systematically over-delete fluent content, revealing a bias toward semantic abstraction over structural fidelity. While fine-tuning achieves SOTA results, it harms generalization. Our findings demonstrate that robustness to speech is shaped by specific training objectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。