arXiv:2604.27403eess.AS2026-04

利用发音方式知识提升电影音轨中被背景音掩盖的语音分离效果

A Knowledge-Driven Approach to Target Speech Extraction in the Presence of Background Sound Effects for Cinematic Audio Source Separation (CASS)

论文配图:A Knowledge-Driven Approach to Target Speech Extraction in the Presence of Background Sound Effects for Cinematic Audio Source Separation (CASS)
图 1 · 摘自论文原文
  • 引入发音方式知识构建特征向量增强语音分离
  • 在复杂背景音中对难分离语音段提升分离效果
  • 适合影视音频后期处理与语音增强场景

我们提出一种基于知识驱动的方法,用于在已录制的电影音频中提取目标语音,尤其针对背景音效干扰的情况。研究的关键知识来源是语音帧中的发音方式,将其转化为知识向量作为特征的一部分,以增强语音分离与目标语音提取能力,因为部分短语音片段常难以与混合背景音分离。在最新的电影音频源分离(CASS)声音解混挑战数据集上的测试表明,使用发音器感知的知识源相比不使用任何知识的方法取得了更优的分离结果,尤其在语音被未指定背景声事件掩盖时表现更佳。

原文摘要 · Abstract (English)

We propose a knowledge-driven approach to speech target extraction in the presence of background sound effects already recorded in cinematic audio. The specific knowledge sources studied are manners of articulation that are detected in speech frames and adopted to form a knowledge vector as a part of features to enhance speech separation and target speech extraction because some short speech segments are often difficult to separate from mixed background sounds. Testing on the recent Sound Demixing Challenge data for cinematic audio source separation (CASS) shows that utilizing articulator-aware knowledge sources produces better separation results than those obtained without using any knowledge, especially for speech segments buried in unspecified background sound events.

语音分离电影音频知识驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。