统一处理多人对话中的说话人识别与验证,提升真实场景下语音分析精度。
Who Spoke When in Multi-Conversation: Target Speaker Tagging Task and Benchmark

- 将说话人分割、验证与识别整合为统一流程,支持长录音中已注册说话人标注。
- 在超150名注册说话人、300个会话的合成数据集上测试,显著优于传统方法。
- 适用于智能会议、语音助手等需精准识别人物身份的场景。
我们提出目标说话人标注(TST)任务,将说话人分割、验证与识别整合为统一工作流,用于多说话人对话场景。给定长录音和预先注册的说话人信息,TST 能检测并标注已知说话人的语音片段,同时拒绝未知说话人。尽管该任务具有重要应用价值,但受限于缺乏合适的评估资源。为此,我们构建了 TST-Bench,一个大规模合成基准数据集,包含超过150名注册说话人、300个持续20至60分钟的会话,以及带全局说话人标签的参考标注。我们定义了涵盖分割与全链路场景的评估协议。在真实与合成数据上的实验表明,TST 提出了现有基准未覆盖的挑战,且专用系统设计相比简单集成现有方案有显著提升。该基准数据集与评估协议已公开发布。
原文摘要 · Abstract (English)
We present target speaker tagging (TST), a task that integrates speaker diarization, verification, and identification into a unified workflow for multi-speaker conversations. Given long recordings and pre-enrolled speakers, TST detects and labels speech segments of known speakers while rejecting unknown ones. Despite its practical importance, research has been limited by the absence of suitable evaluation resources. To address this, we introduce TST-Bench, a large-scale synthetic benchmark with over 150 enrolled speakers, 300 sessions of 20-60 minutes, and reference annotations with global speaker labels. We define an evaluation protocol encompassing diarization and full-pipeline scenarios. Experiments on both real and synthetic data show that TST poses challenges not captured by conventional benchmarks, and that dedicated system design yields significant gains over naive integration of existing solutions. The benchmark dataset and evaluation protocols are publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。