arXiv:2511.12285eess.AScs.CL2025-11中稿 · ICASSP 2026

研究SSL语音模型在低资源下如何捕捉语调,发现其听觉范围受下游任务影响。

How Far Do SSL Speech Models Listen for Tone? Temporal Focus of Tone Representation under Low-resource Transfer

  • 通过探针和梯度分析,考察模型对语调的时序关注范围。
  • 不同语言语调线索持续约100ms(缅泰)或180ms(老越)。
  • 语音识别任务聚焦语言特异性语调,其他任务则过度依赖长时程信息。

声调是众多语言的核心特征,但在自监督学习(SSL)语音模型中仍研究不足,尤其在普通话以外的语言中。本文针对四种具有复杂多样声调系统的语言(缅甸语、泰语、老挝语、越南语),探讨此类模型在多大程度上“倾听”声调,并分析在低资源条件下迁移的表现。作为基线,我们估算出各语言声调线索的时序跨度:缅甸语/泰语约为100ms,老挝语/越南语约为180ms。对微调后的SSL模型进行探针与梯度分析发现,声调迁移效果因下游任务而异:自动语音识别任务的微调使模型关注与语言特定声调线索一致的时长;而韵律和语音相关任务则倾向于过长的关注跨度。结果表明,声调迁移受到下游任务的显著影响,揭示了任务对声调建模中时间焦点的塑造作用。

原文摘要 · Abstract (English)

Lexical tone is central to many languages but remains underexplored in self-supervised learning (SSL) speech models, especially beyond Mandarin. We study four languages with complex and diverse tone systems (Burmese, Thai, Lao, and Vietnamese) to ask how far such models "listen" for tone and how transfer operates in low-resource conditions. As a baseline reference, we estimate the temporal span of tone cues: approximately 100ms (Burmese/Thai) and 180ms (Lao/Vietnamese). Probes and gradient analysis on fine-tuned SSL models reveal that tone transfer varies by downstream task: automatic speech recognition fine-tuning aligns spans with language-specific tone cues, while prosody- and voice-related tasks bias toward overly long spans. These findings indicate that tone transfer is shaped by downstream task, highlighting task effects on temporal focus in tone modeling.

自监督学习语调建模低资源迁移语音分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。