arXiv:2603.26362cs.CV2026-03中稿 · CVPR被引 3

构建手部空间推理基准,揭示大模型在精细手部动作理解上的缺陷。

HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models

  • 基于3D手部数据构建160万道多选题,测试模型对关节角度距离的判断
  • 模型普遍存在幻觉手指、几何误判问题,微调后仍难泛化
  • 注入3D空间知识可零样本提升手势识别准确率10.33%

在机器人辅助手术、芯片制造和虚实交互等高风险场景中,理解人体手部的精细结构至关重要。尽管当前视觉语言模型(VLMs)在通用基准上接近人类表现,但在解析复杂手部姿态的细粒度空间关系方面仍存在明显短板。本文提出HandVQA,一个大规模诊断基准,通过视觉问答评估模型对手部解剖细节的理解。该基准基于高质量3D手部数据集(FreiHAND、InterHand2.6M、FPHA),包含超过160万道受控的多选题,聚焦手部关节间的空间关系,如角度、距离与相对位置。我们评估了多个先进VLMs(LLaVA、DeepSeek、Qwen-VL)在基础与微调设置下的表现,采用轻量级微调方法LoRA。结果揭示模型存在系统性缺陷:包括幻觉手指部件、错误几何解读及泛化能力差。HandVQA不仅暴露这些关键推理缺口,还验证了通过该基准学习的3D空间知识可在零样本条件下迁移,显著提升下游任务性能,如手部手势识别(+10.33%)和手物交互(+2.63%)。

原文摘要 · Abstract (English)

Understanding the fine-grained articulation of human hands is critical in high-stakes settings such as robot-assisted surgery, chip manufacturing, and AR/VR-based human-AI interaction. Despite achieving near-human performance on general vision-language benchmarks, current vision-language models (VLMs) struggle with fine-grained spatial reasoning, especially in interpreting complex and articulated hand poses. We introduce HandVQA, a large-scale diagnostic benchmark designed to evaluate VLMs' understanding of detailed hand anatomy through visual question answering. Built upon high-quality 3D hand datasets (FreiHAND, InterHand2.6M, FPHA), our benchmark includes over 1.6M controlled multiple-choice questions that probe spatial relationships between hand joints, such as angles, distances, and relative positions. We evaluate several state-of-the-art VLMs (LLaVA, DeepSeek and Qwen-VL) in both base and fine-tuned settings, using lightweight fine-tuning via LoRA. Our findings reveal systematic limitations in current models, including hallucinated finger parts, incorrect geometric interpretations, and poor generalization. HandVQA not only exposes these critical reasoning gaps but provides a validated path to improvement. We demonstrate that the 3D-grounded spatial knowledge learned from our benchmark transfers in a zero-shot setting, significantly improving accuracy of model on novel downstream tasks like hand gesture recognition (+10.33%) and hand-object interaction (+2.63%).

视觉语言模型手部识别空间推理3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。