提出上下文感知的查询优化,让声音提取更适应不完整查询。
Context-Aware Query Refinement for Target Sound Extraction: Handling Partially Matched Queries
- 根据音频活动估计自动剔除查询中不存在的声音
- 在部分匹配查询下性能优于传统方法,提升鲁棒性
- 适合真实场景中含冗余信息的语音指令
目标声音提取(TSE)是从音频混合物中提取由查询指定的目标声音。以往研究多集中于完全匹配查询(FMQ)情形,即查询仅包含混合物中实际存在的声音。但在真实场景中,查询可能包含未出现的声音,导致完全不匹配(FUQ)和部分匹配(PMQ)等情形。其中PMQ下的性能下降问题长期被忽视。为提升PMQ条件下的稳健性,本文提出上下文感知查询优化方法:推理时根据估计的声音活动性自动移除查询中不存在的声音类别。实验表明,尽管传统方法在PMQ下性能显著下降,所提方法能有效缓解这一问题,在多种查询条件下均表现优异。
原文摘要 · Abstract (English)
Target sound extraction (TSE) is the task of extracting a target sound specified by a query from an audio mixture. Much prior research has focused on the problem setting under the Fully Matched Query (FMQ) condition, where the query specifies only active sounds present in the mixture. However, in real-world scenarios, queries may include inactive sounds that are not present in the mixture. This leads to scenarios such as the Fully Unmatched Query (FUQ) condition, where only inactive sounds are specified in the query, and the Partially Matched Query (PMQ) condition, where both active and inactive sounds are specified. Among these conditions, the performance degradation under the PMQ condition has been largely overlooked. To achieve robust TSE under the PMQ condition, we propose context-aware query refinement. This method eliminates inactive classes from the query during inference based on the estimated sound class activity. Experimental results demonstrate that while conventional methods suffer from performance degradation under the PMQ condition, the proposed method effectively mitigates this degradation and achieves high robustness under diverse query conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。