arXiv:2605.31589cs.CV2026-05

构建了140万条真实场景手势数据集,用于识别与话语相关的语义手势。

Recognizing Co-Speech Gestures in-the-Wild

论文配图:Recognizing Co-Speech Gestures in-the-Wild
图 1 · 摘自论文原文
  • 提出GRW数据集,涵盖155个词汇的140万条真实视频手势片段。
  • 在17,000个语义手势上实现动作分类、词语识别与时间定位的高精度模型。
  • 适合研究多模态交互、人机理解与真实场景下手势识别的学者使用。

人类在说话时会自然产生手势,但只有少数手势具有视觉表征意义并与特定词语语义相关。本文提出了一个大规模数据集——在野手势识别(Gesture Recognition in the Wild, GRW),包含对应155个词汇的共语手势,共140万条人工标注视频片段,其中17,000个为具有帧级时间边界的语义手势。视频来自公开演讲、访谈和脱口秀等真实场景,覆盖多样说话者与视觉条件。我们还引入视频模型,实现:(a) 手势是否具语义的分类;(b) 手势对应词语的识别;(c) 手势的时间定位。这些模型在GRW数据集上训练与评估,并与多种强基线对比,确立了三项任务的基准结果。数据集、标注与训练模型已公开于项目网站:https://www.robots.ox.ac.uk/~vgg/research/grw。

原文摘要 · Abstract (English)

While humans naturally gesture during speech, only a sparse subset of these co-speech gestures are visually depictive and semantically linked to specific spoken words. In this paper, we introduce a large-scale dataset -- Gesture Recognition in the Wild (GRW), comprising co-speech gestures corresponding to a diverse vocabulary of 155 words. GRW contains 140k manually annotated video clips where the word is spoken, with 17k instances of semantic co-speech gestures including their frame-level temporal boundaries. The video clips are collected 'in the wild' from public-facing discourse, including lectures, talk shows, and interviews, covering a diverse range of speakers and visual conditions. We also introduce video models to: (a) classify gestures as semantic or not; (b) recognize the word corresponding to a co-speech gesture; and (c) temporally localize the gesture. These models are trained and evaluated on the GRW dataset and compared against a range of strong baselines, establishing benchmark results for all three tasks. The dataset, annotations, and trained models are publicly available on the project website: https://www.robots.ox.ac.uk/~vgg/research/grw.

手势识别多模态真实场景数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。