从重叠语音中直接学习持久说话人表征,无需显式跟踪
Unmixing The Crowd: Learning Persistent Speaker Representations from Mixture-Derived Multi-Speaker Embeddings

- 用短重叠片段和排列不变监督训练混合语音表征
- 结合轻量记忆机制实现长时说话人重识别,准确率超90%
- 适用于多种模型架构,适合实时对话系统应用
我们研究是否可直接从局部重叠语音混合中提取持续的对话说话人结构。提出一种教师-学生框架,仅使用短重叠段和排列不变的潜在监督来学习混合派生的多说话人嵌入。尽管未显式训练用于说话人追踪、聚类或对话记忆,但所学嵌入空间在推理时结合轻量级在线记忆机制后,支持长时说话人重识别。此外,所学表征在未见重叠人数条件下仍保持有意义的说话人结构。对比分离优先流程的嵌入,直接从混合中预测的嵌入聚类结构更优。最后,该嵌入在多种架构下对下游目标说话人提取任务均有效。这些发现表明,结合推理时轻量记忆整合,局部混合派生表征可支持持续对话说话人重识别。
原文摘要 · Abstract (English)
We study whether persistent conversational speaker structure can be extracted directly from local overlapping speech mixtures. We propose a teacher-student framework that learns mixture-derived multi-speaker embeddings using only short overlapping segments and permutation-invariant latent supervision. Despite never being explicitly trained for speaker tracking, diarization, or conversational memory, the learned embedding space supports long-form speaker re-identification when combined with a lightweight online memory mechanism during inference. We additionally observe that the learned representation retains meaningful speaker structure under unseen overlap cardinalities. We further show that embeddings extracted from separation-first pipelines exhibit degraded clustering structure compared to embeddings predicted directly from mixtures. Finally, the learned embeddings remain effective for the downstream target speaker extraction task across multiple architectures. These findings suggest that local mixture-derived representations support persistent conversational speaker re-identification when combined with lightweight inference-time memory consolidation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。