arXiv:2409.12415eess.AScs.AI2024-09被引 6

用方向和时间线索,从多通道混响音频中提取目标声音。

Multichannel-to-Multichannel Target Sound Extraction Using Direction and Timestamp Clues

  • 基于时空线索的多通道到多通道语音分离框架
  • 在不同房间环境的多类声音混合信号上表现优异
  • 无需人工设计空间特征,直接处理方向信息

本文提出一种多通道到多通道目标声音提取(M2M-TSE)框架,用于从多通道声音混合信号中分离出特定目标信号。传统目标声音提取通常针对单通道信号,依赖类别标签或时间激活图作为线索。为保留并利用多通道音频中的空间信息,本文提出基于时空线索的多通道信号提取方法,支持方向到达角(DoA)和源激活时间戳等线索。采用基于Transformer的架构,在多种声音类别和不同房间环境合成的多通道信号上成功实现连续分离任务。结果表明,多通道提取任务为深度神经网络引入足够归纳偏置,使其可直接处理方向线索,无需手工构造空间特征。

原文摘要 · Abstract (English)

We propose a multichannel-to-multichannel target sound extraction (M2M-TSE) framework for separating multichannel target signals from a multichannel mixture of sound sources. Target sound extraction (TSE) isolates a specific target signal using user-provided clues, typically focusing on single-channel extraction with class labels or temporal activation maps. However, to preserve and utilize spatial information in multichannel audio signals, it is essential to extract multichannel signals of a target sound source. Moreover, the clue for extraction can also include spatial or temporal cues like direction-of-arrival (DoA) or timestamps of source activation. To address these challenges, we present an M2M framework that extracts a multichannel sound signal based on spatio-temporal clues. We demonstrate that our transformer-based architecture can successively accomplish the M2M-TSE task for multichannel signals synthesized from audio signals of diverse classes in different room environments. Furthermore, we show that the multichannel extraction task introduces sufficient inductive bias in the DNN, allowing it to directly handle DoA clues without utilizing hand-crafted spatial features.

语音分离多通道时空线索Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。