轻量级双阶段音视频目标说话人分离,边端实时运行
Two-stage Audio-Visual Target Speaker Extraction System for Real-Time Processing On Edge Device
- 分两阶段:先用视觉做语音活动检测,再结合音频分离目标声音
- 计算量极低,实测在边缘设备上实时处理无压力
- 适合移动端、嵌入式等算力受限场景的语音增强应用
音视频目标说话人分离(AVTSE)旨在多说话人环境中利用视觉线索分离特定说话人的语音。现有方法通常同时编码视听特征,导致计算复杂度极高,难以在边缘设备上实时运行。为此,我们提出一种两阶段超轻量级AVTSE系统:第一阶段采用紧凑网络仅用视觉信息进行语音活动检测(VAD);第二阶段将VAD结果与音频输入结合,实现目标说话人语音的分离。实验表明,该系统在显著抑制背景噪声和干扰语音的同时,仅消耗极少计算资源,具备边缘设备实时处理能力。
原文摘要 · Abstract (English)
Audio-Visual Target Speaker Extraction (AVTSE) aims to isolate a target speaker's voice in a multi-speaker environment with visual cues as auxiliary. Most of the existing AVTSE methods encode visual and audio features simultaneously, resulting in extremely high computational complexity and making it impractical for real-time processing on edge devices. To tackle this issue, we proposed a two-stage ultra-compact AVTSE system. Specifically, in the first stage, a compact network is employed for voice activity detection (VAD) using visual information. In the second stage, the VAD results are combined with audio inputs to isolate the target speaker's voice. Experiments show that the proposed system effectively suppresses background noise and interfering voices while spending little computational resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。