arXiv:2412.05150cs.CV2024-12被引 7

融合音视频与人体动作信息,提升复杂场景下说话人检测准确率。

BIAS: A Body-based Interpretable Active Speaker Approach

  • 首次结合音频、人脸和身体特征进行说话人检测。
  • 在复杂场景下优于现有方法,尤其在身体信息关键时表现突出。
  • 通过注意力热图实现可解释性,适合需要透明决策的场景。

当前先进的主动说话人检测(ASD)方法主要依赖音频和面部特征,在真实复杂场景中不可持续。尽管这些方法在标准数据集 AVA-ActiveSpeaker 上表现良好,但新发布的更复杂数据集 WASD 显露了其局限性,亟需新方法。为此,我们提出 BIAS,首个同时融合音频、面部与身体信息的模型,以在多变挑战条件下准确预测主动说话人。此外,我们创新性地利用 Squeeze-and-Excitation 模块生成注意力热图并评估特征重要性,实现模型可解释性。为完善可解释性,我们构建了相关动作标注数据集 ASD-Text,微调 ViT-GPT2 生成文本场景描述以辅助理解。实验表明,BIAS 在哥伦比亚开放设置及 WASD 等挑战场景下达到最优性能;在以面部主导的 AVA-ActiveSpeaker 上亦表现竞争力。其可解释性分析揭示不同环境下对 ASD 预测更重要的特征,为可解释性建模提供坚实基线,代码已开源。

原文摘要 · Abstract (English)

State-of-the-art Active Speaker Detection (ASD) approaches heavily rely on audio and facial features to perform, which is not a sustainable approach in wild scenarios. Although these methods achieve good results in the standard AVA-ActiveSpeaker set, a recent wilder ASD dataset (WASD) showed the limitations of such models and raised the need for new approaches. As such, we propose BIAS, a model that, for the first time, combines audio, face, and body information, to accurately predict active speakers in varying/challenging conditions. Additionally, we design BIAS to provide interpretability by proposing a novel use for Squeeze-and-Excitation blocks, namely in attention heatmaps creation and feature importance assessment. For a full interpretability setup, we annotate an ASD-related actions dataset (ASD-Text) to finetune a ViT-GPT2 for text scene description to complement BIAS interpretability. The results show that BIAS is state-of-the-art in challenging conditions where body-based features are of utmost importance (Columbia, open-settings, and WASD), and yields competitive results in AVA-ActiveSpeaker, where face is more influential than body for ASD. BIAS interpretability also shows the features/aspects more relevant towards ASD prediction in varying settings, making it a strong baseline for further developments in interpretable ASD models, and is available at https://github.com/Tiago-Roxo/BIAS.

说话人检测可解释性多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。