真实对话中提取指定说话人语音,挑战真实场景下的语音分离难题。
SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

- 基于真实对话录音,利用目标说话人语音样本进行语音提取。
- 在混叠、回声、噪声等复杂条件下,目标说话人识别准确率提升显著。
- 适合关注实际应用场景的语音处理研究者和工业界开发者。
我们推出了REAL-TSE挑战赛,作为IEEE SLT 2026的卫星挑战赛,聚焦于从真实对话录音中进行目标说话人提取(TSE)。给定多说话人混合音频及一个或多个目标说话人的注册语句,参赛系统需仅恢复目标说话人语音。与模拟朗读语音基准不同,REAL-TSE采用包含自然重叠、混响、噪声、通道失配和对话动态的真实汉语与英语录音。挑战赛设置两个互补赛道:在线赛道用于低延迟流式提取,离线赛道用于全上下文处理。系统评估指标包括词错误率(TER)、说话人相似度(SpkSim)、DNSMOS和目标说话人活动F1。本文概述了任务定义、数据集、基线模型、评估协议、提交系统、条件相关发现以及未来真实世界TSE基准的启示。
原文摘要 · Abstract (English)
We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the target speech. Unlike simulated read-speech benchmarks, REAL-TSE evaluates Mandarin and English recordings that contain natural overlap, reverberation, noise, channel mismatch, and conversational dynamics. The challenge defines two complementary tracks: an Online track for low-latency streaming extraction and an Offline track for full-context processing. Systems are evaluated with Token Error Rate (TER), Speaker Similarity (SpkSim), DNSMOS, and target-speaker activity F1. This overview paper describes the task definition, datasets, baselines, evaluation protocol, submitted systems, condition-wise findings, and lessons for future real-world TSE benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。