拆解语音模型答错原因:决策规则与读出覆盖问题
Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models
- 构建诊断阶梯,分步分析从隐状态到答案的每一步失误
- 平均准确率差27.8点,决策与读出均存在可改进的短板
- 发现模型能理解但未使用的信息,适合模型优化研究者
语音语言模型在副语言任务上的表现常以提示回答的准确率评估,但准确率混合了多个阶段的失败。本文提出生成对齐的诊断阶梯,比较生成答案、选项得分、得分的仿射读出、以及隐藏状态的线性读出在相同答案位置的表现。逐级差异分离出终点、决策规则和读出覆盖的差距。在五个系统与两个情绪语料上,最优解码比生成高出平均27.8个百分点,且所有十种条件下决策规则与读出覆盖差距均为正。无标签得分校正在所有条件下提升生成准确率,表明部分决策规则差距可被修正。在匹配排名的对比中,超出原读出范围的情绪信息能泛化至未见说话人,并在控制声学特征后仍存留,但替换这些外部方向通常对生成答案影响甚微。结果区分了信息可用性与行为使用,定位性能损失发生在决策规则与状态到答案的读出环节。
原文摘要 · Abstract (English)
Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. We introduce a generation-aligned diagnostic ladder that compares the emitted answer, the option logits, an affine readout of those logits, and a linear readout of the hidden state at the same answer token. Successive differences separate endpoint, decision-rule, and readout-coverage gaps. Across five systems and two emotion corpora, state decoding exceeds generation by 27.8 accuracy points on average, and both the decision-rule and readout-coverage gaps are positive in all ten conditions. A label-free logit correction improves generated accuracy in every condition, showing that part of the decision-rule gap is actionable. In rank-matched comparisons, emotion information outside the native readout generalizes to held-out speakers and survives controls for measured acoustic descriptors, but replacing the selected readout-external directions usually has little effect on emitted answers. These results distinguish information availability from behavioral use and localize performance losses across the decision rule and the state-to-answer readout.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。