用大模型语言知识提升音视频语音分离效果,尤其在难场景下表现更好。
ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
- 通过输出约束、中间预测和输入先验三种方式引入大模型语言知识
- 在视觉线索弱、多说话人等挑战场景下显著提升分离准确率
- 适用于语音分离、跨语言语音处理等场景,适合做多模态语音增强的研究者
音视频目标说话人分离(AV-TSE)模型主要依赖目标说话人的视觉线索。然而人类在听觉分离中还会运用语法约束、下一个词预测以及对话背景等语言知识。受此启发,我们提出ELEGANCE框架,通过三种不同引导策略将大语言模型(LLM)的语言知识融入AV-TSE模型:输出语言约束、中间语言预测和输入语言先验。在两种AV-TSE主干网络上,使用RoBERTa、Qwen3-0.6B和Qwen3-4B进行的全面实验验证了该方法的有效性。在视觉线索受损、未见语言、目标说话人切换、干扰说话人增多及域外测试集等挑战性场景中均取得显著性能提升。
原文摘要 · Abstract (English)
Audio-visual target speaker extraction (AV-TSE) models primarily rely on visual cues from the target speaker. However, humans also leverage linguistic knowledge, such as syntactic constraints, next word prediction, and prior knowledge of conversation, to extract target speech. Inspired by this observation, we propose ELEGANCE, a novel framework that incorporates linguistic knowledge from large language models (LLMs) into AV-TSE models through three distinct guidance strategies: output linguistic constraints, intermediate linguistic prediction, and input linguistic prior. Comprehensive experiments with RoBERTa, Qwen3-0.6B, and Qwen3-4B on two AV-TSE backbones demonstrate the effectiveness of our approach. Significant improvements are observed in challenging scenarios, including visual cue impaired, unseen languages, target speaker switches, increased interfering speakers, and out-of-domain test set. Demo page: https://alexwxwu.github.io/ELEGANCE/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。