通过调整填充参数长度,发现大模型能更好适应跨语言语音验证。
An Investigation of Reprogramming for Cross-Language Adaptation in Speaker Verification Systems
- 在语音信号两端添加可学习参数进行对抗重编程
- 参数过多会导致性能饱和甚至下降,大模型上限更高
- 适合需要跨语言适配的语音验证系统研究者
语言不匹配是部署说话人验证(SV)系统中最常见且最具挑战性的域差异之一。对抗重编程在跨语言适配中展现出良好效果,其方法是在输入语音信号两侧添加可学习参数。本文研究了填充参数数量与重编程模型性能之间的关系。在不同规模的SV模型和数据集上进行了充分实验。结果表明,重编程能持续提升跨语言SV性能,但当填充长度过大时,性能会趋于饱和甚至下降。性能主要取决于原始SV模型的容量,而非填充参数数量。规模更大的SV模型具有更高的性能上限,能承受更长的填充而不出现性能退化。
原文摘要 · Abstract (English)
Language mismatch is among the most common and challenging domain mismatches in deploying speaker verification (SV) systems. Adversarial reprogramming has shown promising results in cross-language adaptation for SV. The reprogramming is implemented by padding learnable parameters on the two sides of input speech signals. In this paper, we investigate the relationship between the number of padded parameters and the performance of the reprogrammed models. Sufficient experiments are conducted with different scales of SV models and datasets. The results demonstrate that reprogramming consistently improves the performance of cross-language SV, while the improvement is saturated or even degraded when using larger padding lengths. The performance is mainly determined by the capacity of the original SV models instead of the number of padded parameters. The SV models with larger scales have higher upper bounds in performance and can endure longer padding without performance degradation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。