arXiv:2509.13572cs.RO2025-09被引 2

用视觉语言模型一键识别物体并推断抓握动作,让假手更智能。

Using Visual Language Models to Control Bionic Hands: Assessment of Object Perception and Grasp Inference

  • 单张图像输入,用VLM统一完成物体识别与抓握参数推断
  • 物体名称和形状识别准确率高,但尺寸与手部姿态估计误差较大
  • 适合研究假肢智能控制或视觉语言模型应用的开发者

本研究探讨了利用视觉语言模型(VLMs)提升半自主假手感知能力的潜力。我们提出一个统一基准,评估单一VLM在端到端感知与抓握推理任务中的表现,替代传统需分模块进行目标检测、位姿估计和抓握规划的复杂流程。在包含34张常见物体静态图像的数据集上,测试八种主流VLM,要求其从单张图像中识别物体名称、形状、朝向、尺寸,并推断抓握类型、手腕旋转、手部开合程度及参与手指数量。通过结构化JSON提示词获取输出,分析类别属性准确率、数值估计误差、延迟与成本等指标。结果显示,多数模型在物体名称和形状识别上表现良好,但对尺寸及手部旋转、开合程度的估计差异较大。该工作揭示了VLM作为假肢智能感知模块的当前能力与局限,展示了其在假肢应用中的前景。

原文摘要 · Abstract (English)

This study examines the potential of utilizing Vision Language Models (VLMs) to improve the perceptual capabilities of semi-autonomous prosthetic hands. We introduce a unified benchmark for end-to-end perception and grasp inference, evaluating a single VLM to perform tasks that traditionally require complex pipelines with separate modules for object detection, pose estimation, and grasp planning. To establish the feasibility and current limitations of this approach, we benchmark eight contemporary VLMs on their ability to perform a unified task essential for bionic grasping. From a single static image, they should (1) identify common objects and their key properties (name, shape, orientation, and dimensions), and (2) infer appropriate grasp parameters (grasp type, wrist rotation, hand aperture, and number of fingers). A corresponding prompt requesting a structured JSON output was employed with a dataset of 34 snapshots of common objects. Key performance metrics, including accuracy for categorical attributes (e.g., object name, shape) and errors in numerical estimates (e.g., dimensions, hand aperture), along with latency and cost, were analyzed. The results demonstrated that most models exhibited high performance in object identification and shape recognition, while accuracy in estimating dimensions and inferring optimal grasp parameters, particularly hand rotation and aperture, varied more significantly. This work highlights the current capabilities and limitations of VLMs as advanced perceptual modules for semi-autonomous control of bionic limbs, demonstrating their potential for effective prosthetic applications.

假肢控制视觉语言模型抓握推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。