提出可动态提取维度的说话人嵌入,8维仍保持高识别性能。
M-Vec: Matryoshka Speaker Embeddings with Flexible Dimensions
- 构建分层嵌入结构,按需提取不同精度子维度。
- 在VoxCeleb上实现8维嵌入,验证准确率仍很高。
- 适合资源受限场景下的高效说话人识别应用。
固定维度的说话人嵌入已成为主流方法,通常包含数百至数千维。这些维度是未明确选择的超参数,且无重要性层级。在大规模说话人表示数据库中,降低嵌入维度可显著减少存储和计算成本。然而,直接训练低维表示往往导致性能下降。本文提出马特罗什卡说话人嵌入(Matryoshka Speaker Embeddings),支持从嵌入中动态提取子维度,同时保持性能。该方法在VoxCeleb数据集上验证,证明其可实现极低维度嵌入(如8维),同时维持高水平的说话人验证性能。
原文摘要 · Abstract (English)
Fixed-dimensional speaker embeddings have become the dominant approach in speaker modeling, typically spanning hundreds to thousands of dimensions. These dimensions are hyperparameters that are not specifically picked, nor are they hierarchically ordered in terms of importance. In large-scale speaker representation databases, reducing the dimensionality of embeddings can significantly lower storage and computational costs. However, directly training low-dimensional representations often yields suboptimal performance. In this paper, we introduce the Matryoshka speaker embedding, a method that allows dynamic extraction of sub-dimensions from the embedding while maintaining performance. Our approach is validated on the VoxCeleb dataset, demonstrating that it can achieve extremely low-dimensional embeddings, such as 8 dimensions, while preserving high speaker verification performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。