音频大模型在说话前已暗自推理,且能被读出。
Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace

- 用光谱法读取中间层音频语义,无需思维链。
- 同一音频内容可同时识别多语言概念,38%中文输出。
- 音频信号影响推理,比文本更准确,且提前出现。
音频语言模型像黑箱:我们只能看到输出,看不到内部计算过程。通过在基座模型 Qwen3-Omni 的音频标记位置使用光谱法(logit lens),发现语音问题的答案在输出任何标记前,已在中间层以词语形式变得可读。五项发现:(1) 该读出内容不依赖问题、选项或模型自转录;即使输入无意义文本,仍能重构‘水门事件’、‘总统角色’,最终指向‘尼克松’——隐含多跳推理,无显式思维链。(2) 内容与语言无关:同一音频在多个文字系统中同时浮现,英文输入下38%的最优读出为中文。(3) 包含副语言信息:相同内容,音频输入比纯文本描述更准确,因音频保留声源、说话人角色和情感。(4) 音频信号不在输入端,约在模型深度10%处启动,在中间层(35%-80%)最清晰分离,激活修补实验证明其因果作用早于最后五分之一层。(5) 删除单层可定位功能:音频感知集中于入口层,答案输出集中于输出层,而检索是中间层分布式的。波形替换控制实验(仅声音变化)隔离出音频驱动信号,揭示了音频模型未发声前的深层认知过程。这些是定性洞察,非量化指标。
原文摘要 · Abstract (English)
An audio language model is a black box in a specific way: we see what it says, never what it works out on the way there, and chain-of-thought monitoring helps only if the model writes its reasoning down. Reading a base Qwen3-Omni with a logit lens at the audio-token positions, we find that the answer to a spoken question becomes legible - in words - in the model's middle layers, before it emits any token. Five findings follow. (1) The readout carries concepts in neither the question, the options, nor the model's own transcription: on a clip whose verbatim transcription is empty garbling, it reconstructs Watergate and scandal, passes through the role president, and resolves to Nixon - a hidden multi-hop chain, read with no chain-of-thought. (2) The content is language-agnostic: one audio-inferred concept surfaces in several scripts at once, and 38% of top-1 readouts are Chinese on English inputs. (3) It is paralinguistic: given the same clip as audio and as the model's own emotion-free caption, the audio mind forms the sound source, speaker role, or affect that the caption discards, and answers correctly more often. (4) The audio-driven signal is absent at the input, turns on about a tenth of the way into the network, separates most cleanly from the text prior in the middle band (35-80% of depth), and activation patching shows it is causally used and committed before the last fifth of the layers. (5) Deleting single layers maps the pipeline: reading the sound in is localized to the entry layers and answer delivery to the output layer, while retrieval is distributed across the interior. Throughout, a waveform-swap control - identical text, only the sound changed - isolates the audio-driven signal from a prior over the printed options. This is a qualitative account of what an audio model works out before it speaks: the quantities are controls, not benchmark scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。