arXiv:2603.05813eess.AS2026-03中稿 · Interspeech 2026被引 3

不调模型参数,直接操控激活值来消除语音识别中的口音误差

Activation Steering for Accent Adaptation in Large Audio Language Models

  • 将口音信息视为隐藏表示中的可解释子空间,通过激活方向识别口音差异
  • 在8种口音上实现稳定降错,最高降低12.3%的词错误率
  • 无需微调模型,推理时即可无参数地调整口音适应

口音变异性仍是自动语音识别中的主要错误来源,但多数适配方法依赖参数微调,缺乏对口音信息编码位置的理解。本文将口音变化视为隐藏表示中的可解释子空间,探究是否可在激活空间中直接识别与控制。通过提取分层编码器激活,估计捕捉口音引发表征偏移的均值位移方向。通过向各层注入这些方向并测量其对口音与标准嵌入的对齐效果,构建了分层口音敏感度图谱,发现口音信息集中于中间编码层的狭窄区间。基于此结构,进一步提出无参数口音调控方法,在推理阶段修改表示而不更新模型权重。在8种口音上的实验显示一致的词错误率下降。

原文摘要 · Abstract (English)

Accent variability remains a major source of errors in automatic speech recognition, yet most adaptation methods rely on parameter fine-tuning without understanding where accent information is encoded. We treat accent variation as an interpretable subspace in hidden representations and investigate whether it can be identified and controlled directly in activation space. We extract layer-wise encoder activations and estimate mean-shift directions capturing accent-induced representation shifts. By injecting these directions into individual layers and measuring how they align accented and standard embeddings, we derive a layer-wise accent sensitivity profile, revealing that accent information concentrates in a narrow band of middle encoder layers. Leveraging this structure, we further introduce parameter-free accent steering that modifies representations during inference without updating model weights. Experiments across eight accents show consistent word error rate reductions.

语音识别口音适配无参数调优激活空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。