arXiv:2603.19697eess.AScs.MM2026-03中稿 · Interspeech 2026

将音视频目标说话人分离中的分离与选择解耦,提升语音质量。

Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction

  • 用冻结的纯音频模型做高保真分离,视觉仅负责选目标
  • 引入轻量线性变换矩阵,精准定位目标说话人通道
  • 兼容多种模型架构,适合追求语音自然度的研究者

本文提出一种新的音视频目标说话人提取(AV-TSE)视角,将分离与目标选择过程解耦。传统方法深度融合音视频特征以重学整个分离流程,受限于真实场景数据的噪声,易导致音质天花板。为此,我们提出Plug-and-Steer:将高保真分离任务交由冻结的纯音频骨干网络完成,严格限制视觉模态仅用于目标选择。引入最小化线性变换——潜在引导矩阵(Latent Steering Matrix, LSM),通过重路由骨干网络内部隐状态,将目标说话人锚定至指定通道。在四种代表性架构上的实验表明,该方法有效保留了不同骨干网络的声学先验,感知质量接近原始骨干网络表现。音频样本可访问:https://plugandsteer.github.io

原文摘要 · Abstract (English)

The goal of this paper is to provide a new perspective on audio-visual target speaker extraction (AV-TSE) by decoupling separation and target selection. Conventional AV-TSE systems typically integrate audio and visual features deeply to re-learn the entire separation process, which can act as a fidelity ceiling due to the noisy nature of in-the-wild audio-visual datasets. To address this, we propose Plug-and-Steer, which assigns high-fidelity separation to a frozen audio-only backbone and limits the role of the visual modality strictly to target selection. We introduce the Latent Steering Matrix (LSM), a minimalist linear transformation that re-routes latent features within the backbone to anchor the target speaker to a designated channel. Experiments across four representative architectures show that our method effectively preserves the acoustic priors of diverse backbones, achieving perceptual quality comparable to that of the original backbones. Audio samples are available at: https://plugandsteer.github.io

音视频分离说话人提取模型解耦语音增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。