arXiv:2512.21204cs.CLcs.AI2025-12ACL被引 1

用少样本数据快速适配新语言语音表示,效率比传统方法高100倍。

SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation

  • 将少样本语音表征学习建模为元学习问题,设计双层优化框架。
  • 仅需1小时目标语言音频,语音辨识度和下游任务性能超越现有最佳水平。
  • 适合追求高效语音模型、低资源语言研究的开发者与研究人员。

人类婴儿仅需数百小时语音接触即可掌握新语言的基本单位,远优于当前自监督语音模型的数据饥渴特性。本文提出SpidR-Adapt,通过极少量无标签数据实现对新语言语音单元的快速适应。将该低资源语音表征学习视为元学习问题,构建多任务自适应预训练(MAdaPT)协议,将其形式化为双层优化框架。为实现可扩展的元训练,提出首阶双层优化(FOBLO),避免高昂计算开销。同时通过交替自监督与监督目标进行鲁棒初始化以稳定元训练。实验表明,SpidR-Adapt在少于1小时目标语言音频下训练后,语音辨识度(ABX)及下游口语建模得分(sWUGGY, sBLIMP, tSC)均超越领域内最优结果,相比标准多任务训练提升100倍数据效率。研究揭示了一条面向生物启发、数据高效的通用路径。代码与模型权重已开源:https://github.com/facebookresearch/spidr-adapt。

原文摘要 · Abstract (English)

Human infants, with only a few hundred hours of speech exposure, acquire basic units of new languages, highlighting a striking efficiency gap compared to the data-hungry self-supervised speech models. To address this gap, this paper introduces SpidR-Adapt for rapid adaptation of speech units to new languages using minimal unlabeled data. We cast such low-resource speech representation learning as a meta-learning problem and construct a multi-task adaptive pre-training (MAdaPT) protocol which formulates the adaptation process as a bi-level optimization framework. To enable scalable meta-training under this framework, we propose a novel heuristic solution, first-order bi-level optimization (FOBLO), avoiding heavy computation costs. Finally, we stabilize meta-training by using a robust initialization through interleaved supervision which alternates self-supervised and supervised objectives. Empirically, SpidR-Adapt achieves rapid gains in phonemic discriminability (ABX) and downstream spoken language modeling scores (sWUGGY, sBLIMP, tSC), surpassing in-domain toplines after training on less than 1h of target-language audio and delivering $100\times$ greater data efficiency than standard multi-task training. These findings highlight a practical, architecture-agnostic path toward biologically inspired, data-efficient representations. We open-source the training code and model checkpoints at https://github.com/facebookresearch/spidr-adapt.

语音表征少样本学习元学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。