arXiv:2511.16951cs.CV2025-11被引 2

让模型精准描述手指动作细节,提升人机交互理解能力。

FingerCap: Fine-grained Finger-level Hand Motion Captioning

  • 用关键帧+手部姿态序列恢复细微手指运动时序
  • 在4万条视频数据上显著提升手指级描述准确率
  • 适合手势识别、智能机器人和多模态交互研究者

理解精细的手部动作对视觉感知、具身智能和多模态交流至关重要。本文提出细粒度手指级手部动作描述任务(FingerCap),旨在生成捕捉手指级语义的文本描述。为此,我们构建了包含4万对视频与描述的FingerCap-40K数据集,涵盖简洁指令式手指动作和多样自然的手物交互。为有效评估,采用基于大模型的HandJudge评分体系,衡量手指级准确性与动作完整性。现有Video-MLLMs因时空稀疏性难以捕捉细微高频手指动态。为此,我们提出FiGOP(手指组图像),将每个RGB关键帧配以后续手部关键点直至下一关键帧,轻量级时间编码器将关键点转为运动嵌入并融合至RGB特征。实验表明,主流开闭源Video-MLLM在该任务仍表现不佳,而引入FiGOP的模型在HandJudge与人工评估中均实现稳定提升。

原文摘要 · Abstract (English)

Understanding fine-grained human hand motion is fundamental to visual perception, embodied intelligence, and multimodal communication. In this work, we propose Fine-grained Finger-level Hand Motion Captioning (FingerCap), which aims to generate textual descriptions that capture detailed finger-level semantics of hand actions. To support this task, we curate FingerCap-40K, a large-scale corpus of 40K paired hand-motion videos and captions spanning two complementary sources: concise instruction-style finger motions and diverse, naturalistic hand-object interactions. To enable effective evaluation, we employ HandJudge, a LLM-based rubric that measures finger-level correctness and motion completeness. Temporal sparsity remains a fundamental bottleneck for current Video-MLLMs, since sparse RGB sampling is insufficient to capture the subtle, high-frequency dynamics underlying fine finger motions. As a simple and compute-friendly remedy, we introduce FiGOP (Finger Group-of-Pictures), which pairs each RGB keyframe with subsequent hand keypoints until the next keyframe. A lightweight temporal encoder converts the keypoints into motion embeddings and integrates them with RGB features. FiGOP adapts the classic GOP concept to finger motion, recovering fine temporal cues without increasing RGB density. Experiments on FingerCap-40K show that strong open- and closed-source Video-MLLMs still struggle with finger-level reasoning, while our FiGOP-augmented model yield consistent gains under HandJudge and human studies.

手势识别视频理解多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。