提出数据隐藏评估法,精准衡量语音质量模型真实表现
Unseen but not Unknown: Using Dataset Concealment to Robustly Evaluate Speech Quality Estimation Models
- 引入数据隐藏机制,拆解模型在真实场景中的性能差距
- 使用9个训练集与9个未见数据集验证,显著提升模型泛化能力
- 轻量级对齐器可提升大模型对未知数据的估计效果,适合工业部署
我们提出数据隐藏(DSC),一种严格的语音质量估计模型评估与解释新方法。DSC量化并分解研究结果与实际应用需求之间的性能差距,提供模型行为和数据集特征的上下文信息。通过在多个数据集上训练,结合AlignNet的对齐器(Aligner)缓解语料效应。在九个训练集和九个未见数据集上,用MOSNet、NISQA及基于Wav2Vec2.0的模型进行验证。DSC揭示了模型的泛化能力与局限性,同时允许使用全部可用数据训练。额外结果显示,在9400万参数的Wav2Vec模型中加入仅1000参数的Aligner,显著提升了其对未见数据的语音质量估计能力。
原文摘要 · Abstract (English)
We introduce Dataset Concealment (DSC), a rigorous new procedure for evaluating and interpreting objective speech quality estimation models. DSC quantifies and decomposes the performance gap between research results and real-world application requirements, while offering context and additional insights into model behavior and dataset characteristics. We also show the benefits of addressing the corpus effect by using the dataset Aligner from AlignNet when training models with multiple datasets. We demonstrate DSC and the improvements from the Aligner using nine training datasets and nine unseen datasets with three well-studied models: MOSNet, NISQA, and a Wav2Vec2.0-based model. DSC provides interpretable views of the generalization capabilities and limitations of models, while allowing all available data to be used at training. An additional result is that adding the 1000 parameter dataset Aligner to the 94 million parameter Wav2Vec model during training does significantly improve the resulting model's ability to estimate speech quality for unseen data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。