arXiv:2509.15655cs.CLeess.AS2025-09EMNLP被引 7

揭示语音模型中语法与概念特征的层级编码规律。

Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations

  • 分层分析71项任务中的最小差异对,探测语义与语法特征
  • 所有模型在语法特征编码上均优于概念特征
  • 首次系统验证语音模型的上下文语义层次结构

基于Transformer的语音语言模型(SLMs)显著提升了神经语音识别与理解能力。尽管已有研究考察了SLMs对浅层声学和音素特征的编码能力,但其对复杂句法与概念特征的建模程度仍不明确。受大语言模型语言能力评估的启发,本研究首次系统评估了自监督学习(S3M)、自动语音识别(ASR)、语音压缩(codec)以及作为听觉大语言模型(AudioLLMs)编码器的各类SLMs中,上下文相关的句法与语义特征的存在性。通过跨71项任务的最小差异对设计与诊断特征分析,我们的分层与时间解析分析发现:所有语音模型在语法特征上的编码比概念特征更稳健。

原文摘要 · Abstract (English)

Transformer-based speech language models (SLMs) have significantly improved neural speech recognition and understanding. While existing research has examined how well SLMs encode shallow acoustic and phonetic features, the extent to which SLMs encode nuanced syntactic and conceptual features remains unclear. By drawing parallels with linguistic competence assessments for large language models, this study is the first to systematically evaluate the presence of contextual syntactic and semantic features across SLMs for self-supervised learning (S3M), automatic speech recognition (ASR), speech compression (codec), and as the encoder for auditory large language models (AudioLLMs). Through minimal pair designs and diagnostic feature analysis across 71 tasks spanning diverse linguistic levels, our layer-wise and time-resolved analysis uncovers that 1) all speech encode grammatical features more robustly than conceptual ones.

语音模型句法分析层级表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。