让语音识别适应不同长度,短语音也能精准辨人
DAME: Duration-Aware Matryoshka Embedding for Duration-Robust Speaker Verification
- 按语音时长分级生成嵌入,短句用低维特征,长句用高维细节
- 在1秒短语音上等错误率显著降低,长语音性能不降反升
- 兼容各种模型架构,训练和微调都可直接替换现有方案
短语音说话人验证因语音片段中区分性线索有限而极具挑战。现有方法虽聚焦增强说话人编码器,但嵌入学习策略仍强制使用单一固定维度表示,无法匹配不同长度语音的信息量。本文提出持续感知的套娃嵌入(DAME),一种与模型无关的框架,构建与语音时长对齐的嵌套子嵌入层级:低维表示从短语音中捕捉紧凑的说话人特征,高维则编码长语音中的丰富细节。DAME支持从头训练和微调,可直接替代传统大间隔微调,且在各类时长下均表现更优。在VoxCeleb1-O/E/H和VOiCES评测集上,DAME在1秒及其它短时长测试中持续降低等错误率,同时保持全时长性能且无额外推理开销。该提升在多种说话人编码器架构下,无论通用训练或微调设置,均具泛化能力。
原文摘要 · Abstract (English)
Short-utterance speaker verification remains challenging due to limited speaker-discriminative cues in short speech segments. While existing methods focus on enhancing speaker encoders, the embedding learning strategy still forces a single fixed-dimensional representation reused for utterances of any length, leaving capacity misaligned with the information available at different durations. We propose Duration-Aware Matryoshka Embedding (DAME), a model-agnostic framework that builds a nested hierarchy of sub-embeddings aligned to utterance durations: lower-dimensional representations capture compact speaker traits from short utterances, while higher dimensions encode richer details from longer speech. DAME supports both training from scratch and fine-tuning, and serves as a direct alternative to conventional large-margin fine-tuning, consistently improving performance across durations. On the VoxCeleb1-O/E/H and VOiCES evaluation sets, DAME consistently reduces the equal error rate on 1-s and other short-duration trials, while maintaining full-length performance with no additional inference cost. These gains generalize across various speaker encoder architectures under both general training and fine-tuning setups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。