语音持续学习应聚焦表示结构演化,而非孤立任务记忆。
Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems

- 从表示几何角度重新定义语音持续学习
- 揭示当前假设与语音大模型行为的不匹配
- 为研究者提供新方向与开放问题
语音与音频系统运行于本质非平稳的环境中,但该领域的持续学习(CL)研究,尤其是在基础模型时代,仍显碎片化,未能充分考虑声学表示的耦合性与几何敏感性。现代语音基础模型在共享潜在空间中高度纠缠地联合编码语言、说话人及副语言信息。因此,持续学习本质上是关于保持并演化共享表示结构,而非保留孤立的任务知识。本文从表示中心视角重新审视语音持续学习,提出一种新分类法,依据非平稳声学条件下底层表示几何如何演变来组织持续学习方法。我们进一步识别出现有持续学习假设与语音基础模型行为之间的关键错配,并提出一系列开放挑战与未来研究方向。
原文摘要 · Abstract (English)
Speech and audio systems operate in inherently non-stationary environments, yet continual learning (CL) research in this domain, especially in the foundation model era, remains fragmented that fail to account for the coupled, geometry-sensitive nature of acoustic representations. Modern speech foundation models operate over highly entangled, continuous representations that jointly encode linguistic, speaker, and paralinguistic factors within a shared latent space. CL is therefore fundamentally about preserving and evolving shared representation structure rather than retaining isolated task knowledge. In this work, we revisit CL for speech from a representation-centered perspective, and introduce a new taxonomy that organizes CL according to how underlying representation geometry evolves under non-stationary acoustic conditions. We further identify key mismatches between current CL assumptions and speech foundation model behavior, and finally outline a set of open challenges and future research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。