统一框架实现多种线索下的目标说话人分离,灵活适配真实场景。
WeSep: A Modular and Cue-Composable Framework for Target Speaker Extraction

- 将不同模态线索解耦接入,支持动态组合与配置
- 在多种线索下均保持稳定优化,适应实际环境变化
- 适合需要多线索融合的语音分离任务研究者
目标说话人提取(TSE)旨在给定辅助线索时从重叠语音中分离出指定说话人。现有系统通常针对特定线索类型设计,当线索可用性变化时灵活性不足。本文提出WeSep,一个统一框架,将TSE重构为异质线索条件学习问题。在WeSep中,线索模块与分离主干通过标准化接口解耦,支持可配置的线索注入和多样模态的灵活集成。该设计使得在同一优化框架内可系统研究线索结构、模态内与跨模态交互以及动态线索可用性,提升对真实场景的适应能力。在注册、空间、视觉和文本线索上的实验揭示了模态依赖特性,并验证了在异质线索条件下的稳定优化。工具包将公开可用。
原文摘要 · Abstract (English)
The study of Target Speaker Extraction (TSE) aims to isolate a desired speaker from overlapping speech mixture given auxiliary cues. Existing systems are typically designed for specific cue types, limiting flexibility when cue availability varies across scenarios. We present WeSep, a unified framework that reformulates TSE as a heterogeneous cue-conditioned learning problem. In WeSep, cue modules and separator backbones are decoupled through standardized interfaces, enabling configurable cue injection and flexible integration of diverse modalities. The design enables systematic study of cue structure, intra- and cross-modal interaction, and dynamic cue availability within a shared optimization framework, facilitating adaptation to real-world conditions. Experiments across enrollment, spatial, visual, and textual cues reveal modality-dependent characteristics and demonstrate stable optimization under heterogeneous cue availability. The toolkit will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。