模型能听懂语气却不说出来,关键在如何使用已有的语音信息。
Represented but Ignored: A Causal Account of Prosodic Underuse in Audio-Language Models

- 用分阶段探针定位模型在语音理解中的失败环节。
- 语音语调信息在模型中被保留但未充分表达于最终输出。
- 仅需修改特定层隐藏状态即可部分恢复正确语气判断。
人类语音富含韵律,承载着词汇之外的语言与情感信息。理想的大型音频-语言模型(audio-LLM)应不仅能转录说了什么,还能理解怎么说。然而,仅靠行为评估无法揭示模型为何在韵律输入上表现不佳——错误可能源于声学信息丢失、内部解释错误,或未能利用模型内已存在的表示。本文提出一种分阶段探针梯度,用于定位音频-语言模型中的故障模式。在四个仅理解类音频-语言模型中,韵律信息通常在音频路径中得以保留,并可在大语言模型的后期状态中解码。然而,这些信息仅部分体现在模型最终输出中。通过针对性的隐藏状态干预测试其因果性,发现每项干预均使答案分布朝预期方向移动;在大多数模型-任务组合中,单次在相关层的编辑即可引导模型向被抑制的韵律决策靠近,尽管这种恢复是方向性的而非对正确类别的精确还原。特征级分析进一步表明,该可恢复信号可通过一个小子空间表达,其中部分高重要性特征与已知携带韵律信息的声学线索一致。在匹配内容的对比实验中,结果将反复出现的瓶颈定位为‘使用’而非‘感知’韵律。模型即便听懂并正确表征了韵律线索,仍可能在回答中忽略其表达。
原文摘要 · Abstract (English)
Human speech is richly expressive, with prosody carrying linguistic and emotional information beyond the lexical content. A capable large audio-language model (audio-LLM) should therefore support expressive speech understanding, not only transcribing what was said but also interpreting how it was said. Yet behavioral evaluations alone cannot reveal why a model fails on prosodic input. An error may reflect loss of acoustic information, incorrect internal interpretation, or failure to use a representation that is already available inside the model. We introduce a stage-specific probe ladder for localizing these failure modes in audio-LLMs. Across four understanding-only audio-LLMs, prosodic information is usually preserved in the audio path and decodable in late LLM states. Yet it is only partially expressed in the model's final response. We test the causal status of this latent representation with targeted hidden-state interventions. Every intervention shifts the answer distribution in the predicted direction, and in most model--task cells a single edit at the relevant layer is sufficient to drive the model toward the suppressed prosodic decision, though this recovery is directional rather than a selective restoration of the correct class. Feature-level analysis further suggests that this recoverable signal can be expressed through a small subspace. Some of the highest-attribution features in this analysis align with acoustic cues known to carry prosodic information. Within the matched-content contrasts we test, these results locate the recurring bottleneck not in perceiving prosody but in using it. Models that hear and correctly represent a prosodic cue can still fail to express it in their answers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。