arXiv:2510.00952eess.AScs.SD2025-10被引 3

CL-UZH在2024年说话人识别挑战赛中提交了音视频双模系统,提升识别准确率。

CL-UZH submission to the NIST SRE 2024 Speaker Recognition Evaluation

  • 使用Kaldi训练的X-vector系统处理纯音频数据,结合视觉模型实现音视频融合。
  • 基于VoxBlink2和VoxCeleb2预训练模型,在闭集与开集条件下均取得良好性能。
  • 针对闭集任务重新训练模型,使用CTS超集数据集,适配实际应用场景。

CL-UZH团队为NIST SRE 2024挑战赛的固定条件和开放条件各提交一个系统。在闭集条件下,纯音频试验采用基于Kaldi开发的X-vector系统;音视频结果仅使用专为视觉模态训练的模型。针对开集和闭集条件,分别提交两组结果,其中一组基于在VoxBlink2和VoxCeleb2数据集上预训练的模型。此外,为闭集任务从头训练了一个基于CTS超集数据集的X-vector模型。本报告还详细分析了所提系统在SRE24评估中的表现。

原文摘要 · Abstract (English)

The CL-UZH team submitted one system each for the fixed and open conditions of the NIST SRE 2024 challenge. For the closed-set condition, results for the audio-only trials were achieved using the X-vector system developed with Kaldi. For the audio-visual results we used only models developed for the visual modality. Two sets of results were submitted for the open-set and closed-set conditions, one based on a pretrained model using the VoxBlink2 and VoxCeleb2 datasets. An Xvector-based model was trained from scratch using the CTS superset dataset for the closed set. In addition to the submission of the results of the SRE24 evaluation to the competition website, we talked about the performance of the proposed systems on the SRE24 evaluation in this report.

说话人识别音视频融合X-vector

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。