arXiv:2504.02244cs.CV2025-04CVPR被引 15

首个面向多人手势交互的大规模数据集,助力理解社交场景中的自然手势。

SocialGesture: Delving into Multi-person Gesture Understanding

  • 构建首个聚焦多人互动的手势数据集,覆盖多样真实场景。
  • 支持视频识别与时间定位等多任务,推动社交手势研究。
  • 提出视觉问答新任务,评估模型在社交语境下的理解能力。

现有手势识别研究多忽视多人互动,而此类交互对理解自然手势的社交语境至关重要。当前数据集的局限性导致手势与语言、语音等模态难以对齐。为此,我们提出 SocialGesture,首个专为多人手势分析设计的大规模数据集。该数据集涵盖丰富自然场景,支持视频识别、时间定位等多项任务,为复杂社交互动中的手势研究提供重要资源。此外,我们引入一种新型视觉问答(VQA)任务,用于评估视觉语言模型(VLMs)在社交手势理解上的表现。研究发现当前手势模型存在显著局限,为未来改进方向提供启示。SocialGesture 数据集已开源,可于 huggingface.co/datasets/IrohXu/SocialGesture 获取。

原文摘要 · Abstract (English)

Previous research in human gesture recognition has largely overlooked multi-person interactions, which are crucial for understanding the social context of naturally occurring gestures. This limitation in existing datasets presents a significant challenge in aligning human gestures with other modalities like language and speech. To address this issue, we introduce SocialGesture, the first large-scale dataset specifically designed for multi-person gesture analysis. SocialGesture features a diverse range of natural scenarios and supports multiple gesture analysis tasks, including video-based recognition and temporal localization, providing a valuable resource for advancing the study of gesture during complex social interactions. Furthermore, we propose a novel visual question answering (VQA) task to benchmark vision language models'(VLMs) performance on social gesture understanding. Our findings highlight several limitations of current gesture recognition models, offering insights into future directions for improvement in this field. SocialGesture is available at huggingface.co/datasets/IrohXu/SocialGesture.

手势识别多人交互数据集视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。