针对真实对话场景的语音分离,提出新系统并取得领先表现。
The SonicAGI System for the REAL-TSE Challenge
- 用仿真数据混合真实会议重叠,结合去噪镜像辅助训练。
- 在线版延迟仅96毫秒,离线版在音质与说话人保真间平衡优化。
- 适合需要低延迟或高保真的实际语音分离应用。
真实世界目标说话人提取(TSE)仍具挑战性,因目标语音、干扰信号与注册音频在混响、噪声及不规则对话重叠条件下录制,存在声学条件不匹配问题。本文介绍SonicAGI参与IEEE SLT 2026 REAL-TSE挑战的方案。我们采用数据驱动方法,将纯净语音的全仿真混合与真实会议重叠相结合,并使用冻结的离线增强器提供真实目标的去噪镜像,用于辅助监督。在线赛道引入SwiftNet-Lookahead,其在严格因果迭代分离器前加入单个有限前瞻模块,保持总系统延迟为96毫秒。离线赛道采用帧级注册交叉注意力的USEF-TFGridNet,并加入幅度域融合阶段,在感知质量与说话人保真度之间实现权衡。官方评估中,SwiftNet-Lookahead在第一赛道排名第二,USEF-TFGridNet在第二赛道排名第五,均超过挑战基线。结果表明,面向真实数据的训练与赛道特化建模对对话式TSE有效。
原文摘要 · Abstract (English)
Real-world target speaker extraction (TSE) remains challenging because target speech, interference, and enrollment are recorded under mismatched acoustic conditions with reverberation, noise, and irregular conversational overlap. This paper describes the SonicAGI submission to the REAL-TSE Challenge (IEEE SLT 2026). We take a data-centric approach that combines fully simulated mixtures from clean speech with real meeting overlaps, and use a frozen offline enhancer to provide a denoised mirror of real targets for auxiliary supervision. For the online track, we introduce SwiftNet-Lookahead, which inserts a single bounded-lookahead module before a strictly causal iterative separator and keeps the total system latency at 96 ms. For the offline track, we use a frame-level enrollment cross-attention USEF-TFGridNet with a magnitude-domain fusion stage that trades off perceptual quality and speaker fidelity. In the official evaluation, SwiftNet-Lookahead ranks second in Track~1 and USEF-TFGridNet ranks fifth in Track~2, both exceeding the challenge baselines. These results suggest that real-data-oriented training and track-specific modeling are effective for conversational TSE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。