arXiv:2606.12199eess.AScs.CL2026-06中稿 · Interspeech 2026 l…

找对语音粒度,让语音模型更懂文字逻辑。

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation

论文配图:Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation
图 1 · 摘自论文原文
  • 用固定信息率测试不同语音帧率,选最优语义对齐点。
  • 4.17赫兹时语音问答表现最佳,优于高或低帧率。
  • 适合研究语音与文本对齐、多模态推理的开发者。

语音对话模型通常基于文本大模型构建,但以语音为条件时推理能力下降。我们发现部分原因是时间粒度不匹配:语音标记在语义一致下比文本冗长得多,稀释了每标记的语义密度,削弱了原本面向文本的推理机制。本文将语音标记设计视为表示选择问题,在固定信息率下,使用冻结的LLM主干网络测试不同帧率。为实现低帧率可行性,引入因子化FSQ和轻量级非自回归音频语言模型头,使容量达到近300比特/帧且保持高效预测。移除瓶颈后,测试从50赫兹到2.08赫兹的帧率及对齐深度,发现语音问答在4.17赫兹处表现最佳,且需中间层表示对齐。

原文摘要 · Abstract (English)

Spoken dialogue models typically start from text LLM backbones, yet reasoning often degrades when conditioning on speech instead of text. We attribute part of this modality gap to a temporal-granularity mismatch: speech tokens are temporally redundant and far longer than text under matched semantics, diluting per-token semantic density and weakening text-native reasoning dynamics. We study speech token design as a representation selection problem and sweep frame rates under a frozen LLM backbone with a fixed information rate. To make low frame rates feasible, we introduce factorized FSQ and a lightweight non-autoregressive audio LM head, scaling capacity to nearly 300\,bits/frame without sacrificing efficient prediction. With the bottleneck removed, we sweep frame rates (50$\rightarrow$2.08\,Hz) and alignment depth, and observe a consistent best regime for speech QA at 4.17\,Hz with intermediate-layer representation alignment.

语音理解多模态对齐大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。