arXiv:2508.15316cs.CLcs.LG2025-08中稿 · 8th International …被引 3

用120毫秒窗口实现跨语言通用音素编码,无需上下文

CUPE: Contextless Universal Phoneme Encoder for Language-Agnostic Speech Processing

  • 基于音素级时长窗口的独立处理,轻量化设计
  • 在多语言数据上达到与主流方法相当的跨语言性能
  • 适合需要纯净音素表示的语音任务,如跨语言语音识别

通用音素识别通常需分析长语音段和语言特定模式。许多语音处理任务需要不受上下文影响的纯音素表示,这促使我们开发了CUPE——一种轻量级模型,仅需120毫秒(约一个音素长度)即可捕捉关键音素特征。CUPE独立处理短而固定的时窗,尽管参数量少于现有方法,仍通过学习所有语言共有的基础声学模式,实现了具有竞争力的跨语言表现。我们在多种语言上进行了监督与自监督训练,并在UCLA语音语料库上进行零样本测试,结果表明其具备强大的跨语言泛化能力,证明通过建模音素级时窗内的基本声学模式,可实现有效的通用语音处理。

原文摘要 · Abstract (English)

Universal phoneme recognition typically requires analyzing long speech segments and language-specific patterns. Many speech processing tasks require pure phoneme representations free from contextual influence, which motivated our development of CUPE - a lightweight model that captures key phoneme features in just 120 milliseconds, about one phoneme's length. CUPE processes short, fixed-width windows independently and, despite fewer parameters than current approaches, achieves competitive cross-lingual performance by learning fundamental acoustic patterns common to all languages. Our extensive evaluation through supervised and self-supervised training on diverse languages, including zero-shot tests on the UCLA Phonetic Corpus, demonstrates strong cross-lingual generalization and reveals that effective universal speech processing is possible through modeling basic acoustic patterns within phoneme-length windows.

音素编码跨语言轻量化语音处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。