用嵌入有效秩统一衡量音频模型性能,揭示其缩放规律
Unify Variables in Neural Scaling Laws for General Audio Representations via Embedding Effective Rank
- 以嵌入有效秩为统一指标,整合长度、维度、深度等多变量影响
- 发现嵌入秩与表示质量呈幂律关系,可预测模型性能
- 为音频基础模型的规模扩展提供理论与实证指导
缩放定律深刻影响了计算机视觉和自然语言处理中模型性能的理解,但在通用音频表征学习中的应用仍不充分。核心挑战在于音频表征质量受音频长度、嵌入维度、模型深度、架构、数据量等多种变量共同影响,且这些变量难以分离或解析表达。本文通过引入嵌入有效秩(RankMe)作为统一度量,系统研究通用音频表征的缩放规律。RankMe实现了无标签、信息论意义上的音频嵌入量化,使我们能在广泛超参数空间(包括模型大小、训练数据量、计算预算、架构配置等)中分析缩放行为。实验结果表明,RankMe与表示质量间存在稳定的幂律关系,表明嵌入有效秩可作为评估和预测音频表征性能的可靠代理。本工作不仅验证了经典缩放原则在通用音频领域的适用性,还为未来音频基础模型的规模扩展策略提供了理论扎实、实证稳健的框架。
原文摘要 · Abstract (English)
Scaling laws have profoundly shaped our understanding of model performance in computer vision and natural language processing, yet their application to general audio representation learning remains underexplored. A key challenge lies in the multifactorial nature of general audio representation-representation quality is jointly influenced by variables such as audio length, embedding dimensionality, model depth, model architecture, data volume, etc., many of which are difficult to isolate or express analytically. In this work, we present a systematic study of scaling laws for general audio representations by utilizing embedding effective rank (RankMe) as a unifying metric that encapsulates the impact of diverse variables on representation quality. RankMe enables a label-free, information-theoretic quantification of audio embeddings, allowing us to examine scaling behaviors across a wide hyper-parameter space, including model size, training data volume, computational budget, architectural configurations, etc. Our empirical findings reveal a consistent power-law relationship between RankMe and representation quality, suggesting that embedding effective rank serves as a reliable proxy for assessing and predicting model performance in audio representation learning. This work not only validates the applicability of classical scaling principles to the general audio domain but also offers a theoretically grounded and empirically robust framework for guiding future model scaling strategies in audio foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。