arXiv:2604.19635cs.SDcs.AI2026-04被引 2

首个面向实时目标说话人分离的自回归模型,解决延迟与性能失衡问题。

StarTSE: Towards Streaming Target Speaker Extraction via Chunk-wise Interleaved Splicing of Autoregressive Language Model

论文配图:StarTSE: Towards Streaming Target Speaker Extraction via Chunk-wise Interleaved Splicing of Autoregressive Language Model
图 1 · 摘自论文原文
  • 分块交错拼接机制实现高效稳定流式推理
  • 低延迟下保持100%稳定性与更优语音可懂度
  • 适合实时语音分离场景,尤其对硬件要求低

生成模型虽在目标说话人分离(TSE)上取得新突破,但其对全局上下文的依赖限制了实时应用。直接适配流式场景常因训练与推理不匹配导致性能灾难性下降。为此,我们提出首个专为流式TSE设计的自回归(AR)模型。方法引入分块交错拼接范式,确保高效稳定流式推理;并通过历史上下文优化机制,利用历史信息缓解语音片段间的边界不连续。在Libri2Mix数据集上的实验表明,尽管自回归基线在低延迟下性能下降,我们的方法仍保持100%稳定性且语音可懂度更优;流式结果甚至媲美或超越离线基线。此外,模型在消费级GPU上达到0.248的实时因子(RTF)。本工作实证表明,通过分块交错拼接范式,自回归生成主干可适用于低延迟敏感应用。

原文摘要 · Abstract (English)

While generative models have set new benchmarks for Target Speaker Extraction (TSE), their inherent reliance on global context precludes deployment in real-time applications. Direct adaptation to streaming scenarios often leads to catastrophic inference performance degradation due to the severe mismatch between training and streaming inference. To bridge this gap, we present the first autoregressive (AR) models tailored for streaming TSE. Our approach introduces a Chunk-wise Interleaved Splicing Paradigm that ensures highly efficient and stable streaming inference. To ensure the coherence between the extracted speech segments, we design a historical context refinement mechanism that mitigates boundary discontinuities by leveraging historical information. Experiments on Libri2Mix show that while AR generative baseline exhibits performance degradation at low latencies, our approach maintains 100% stability and superior intelligibility. Furthermore, our streaming results are comparable to or even surpass offline baselines. Additionally, our model achieves a Real-Time-Factor (RTF) of 0.248 on consumer-level GPUs. This work provides empirical evidence that AR generative backbones are viable for latency-sensitive applications through the Chunk-wise Interleaved Splicing Paradigm.

语音分离流式处理自回归模型实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。