arXiv:2601.12205cs.SDcs.AI2026-01被引 2

神经音频编码器能跨语言通用,但仅用语音训练会弱化非语音任务表现。

Do Neural Codecs Generalize? A Controlled Study Across Unseen Languages and Non-Speech Tasks

  • 从零训练编码器,严格控制配置实现公平对比。
  • 语音模型在环境音、音乐等任务上性能下降,但加入非语音数据可提升其表现。
  • 适合研究多任务音频建模或希望提升编码器泛化能力的开发者。

本文研究了神经音频编码器(NACs)三个关键但未被充分探索的泛化特性:(i) NACs 在预训练阶段能否泛化到未见过的语言;(ii) 仅用语音训练的 NACs 能否有效应用于环境声、音乐和动物叫声等非语音任务;(iii) 在预训练中加入非语音数据是否能同时提升语音与非语音任务的表现。现有研究多依赖现成 NACs 比较,受限于实现差异。本工作采用严格控制的配置,从头训练 NACs 并精心设计预训练数据,确保公平比较。通过 11 项指标全面评估信号重建质量与下游应用表现。结果表明:NACs 可在预训练中泛化至未见语言;仅语音训练的 NACs 在非语音任务上表现下降;在预训练中引入非语音数据能显著提升非语音任务性能,同时保持与原版相当的语音任务表现。

原文摘要 · Abstract (English)

This paper investigates three crucial yet underexplored aspects of the generalization capabilities of neural audio codecs (NACs): (i) whether NACs can generalize to unseen languages during pre-training, (ii) whether speech-only pre-trained NACs can effectively generalize to non-speech applications such as environmental sounds, music, and animal vocalizations, and (iii) whether incorporating non-speech data during pre-training can improve performance on both speech and non-speech tasks. Existing studies typically rely on off-the-shelf NACs for comparison, which limits insight due to variations in implementation. In this work, we train NACs from scratch using strictly controlled configurations and carefully curated pre-training data to enable fair comparisons. We conduct a comprehensive evaluation of NAC performance on both signal reconstruction quality and downstream applications using 11 metrics. Our results show that NACs can generalize to unseen languages during pre-training, speech-only pre-trained NACs exhibit degraded performance on non-speech tasks, and incorporating non-speech data during pre-training improves performance on non-speech tasks while maintaining comparable performance on speech tasks.

音频编码泛化能力多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。