arXiv:2503.22668cs.CV2025-03ICCV被引 6

弱监督学习视频中手势与语音文本的关联,提升跨模态理解能力。

Understanding Co-speech Gestures in-the-wild

  • 构建三模态联合表示,通过对比学习和耦合损失实现弱监督训练。
  • 在野外视频上超越大模型,在手势检索、定位和说话人检测任务上表现优异。
  • 揭示语音与文本提供不同手势信号,适合多模态交互与人机协作研究者。

共言语手势在非语言交流中至关重要。本文提出一种面向野外场景的共言语手势理解新框架,设计三项新任务与基准:(i)基于手势的检索,(ii)手势词定位,(iii)利用手势进行主动说话人检测。提出一种三模态(视频-手势-语音-文本)联合表示学习方法,结合全局短语对比损失与局部手势-词耦合损失,实现了从野外视频中弱监督学习强手势表征。实验表明,所学表征优于先前方法,包括大型视觉-语言模型。进一步分析显示,语音与文本模态捕捉到不同的手势相关信号,验证了共享三模态嵌入空间的优势。数据集、模型与代码已公开。

原文摘要 · Abstract (English)

Co-speech gestures play a vital role in non-verbal communication. In this paper, we introduce a new framework for co-speech gesture understanding in the wild. Specifically, we propose three new tasks and benchmarks to evaluate a model's capability to comprehend gesture-speech-text associations: (i) gesture based retrieval, (ii) gesture word spotting, and (iii) active speaker detection using gestures. We present a new approach that learns a tri-modal video-gesture-speech-text representation to solve these tasks. By leveraging a combination of global phrase contrastive loss and local gesture-word coupling loss, we demonstrate that a strong gesture representation can be learned in a weakly supervised manner from videos in the wild. Our learned representations outperform previous methods, including large vision-language models (VLMs). Further analysis reveals that speech and text modalities capture distinct gesture related signals, underscoring the advantages of learning a shared tri-modal embedding space. The dataset, model, and code are available at: https://www.robots.ox.ac.uk/~vgg/research/jegal.

手势理解多模态弱监督视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。