用视觉引导音频定位,提升复杂场景下多人说话者追踪精度
STNet: Deep Audio-Visual Fusion Network for Robust Speaker Tracking
- 视觉引导构建增强声场图,统一声像定位空间
- 跨模态注意力融合多模态特征,提升交互建模能力
- 质量感知模块增强鲁棒性,适合复杂多说话者场景
音频-视觉说话者追踪旨在利用多传感器平台采集的信号确定场景中人目标的位置,其准确性和鲁棒性可通过多模态融合方法提升。尽管已有多种融合方法被提出以建模多模态间的相关性,但对说话者追踪问题中音频与视觉信号间跨模态交互的挖掘仍不充分。为此,本文提出一种新型说话者追踪网络(STNet),包含深度音视频融合模型。设计了视觉引导的声学测量方法,在统一定位空间中融合异构线索,利用相机模型将视觉观测结果用于构建增强声场图。特征融合采用跨模态注意力模块,联合建模多模态上下文与交互关系,进一步强化音视频特征间的相关性。此外,基于质量感知模块的STNet追踪器可处理多说话者场景,通过评估多模态观测的可靠性,实现在复杂环境中的鲁棒追踪。在AV16.3和CAV3D数据集上的实验表明,所提出的STNet追踪器优于单模态方法及当前最先进的音视频说话者追踪方法。
原文摘要 · Abstract (English)
Audio-visual speaker tracking aims to determine the location of human targets in a scene using signals captured by a multi-sensor platform, whose accuracy and robustness can be improved by multi-modal fusion methods. Recently, several fusion methods have been proposed to model the correlation in multiple modalities. However, for the speaker tracking problem, the cross-modal interaction between audio and visual signals hasn't been well exploited. To this end, we present a novel Speaker Tracking Network (STNet) with a deep audio-visual fusion model in this work. We design a visual-guided acoustic measurement method to fuse heterogeneous cues in a unified localization space, which employs visual observations via a camera model to construct the enhanced acoustic map. For feature fusion, a cross-modal attention module is adopted to jointly model multi-modal contexts and interactions. The correlated information between audio and visual features is further interacted in the fusion model. Moreover, the STNet-based tracker is applied to multi-speaker cases by a quality-aware module, which evaluates the reliability of multi-modal observations to achieve robust tracking in complex scenarios. Experiments on the AV16.3 and CAV3D datasets show that the proposed STNet-based tracker outperforms uni-modal methods and state-of-the-art audio-visual speaker trackers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。