让假肢手学会抓取没见过的物体,提升残障人士生活独立性。
Grasp-HGN: Grasping the Unexpected
- 用视觉语言模型模拟人类推理,根据物体特征推断合适抓法。
- 对未见过的物体抓取准确率达50.2%,比现有模型高13.5个百分点。
- 边缘与云端协同部署,既快又准,适合真实环境中的假肢控制。
对于桡骨截肢者而言,机器人假肢手有望恢复日常活动能力。为推动下一代假肢手控制设计,亟需解决模型在实验室外干扰下的鲁棒性不足及对新环境泛化能力差的问题。由于现有数据集所含交互对象数量有限,而现实世界中物体种类几乎无穷,当前抓取模型在未见物体上表现不佳,影响用户独立性和生活质量。为此:(i)我们定义了语义投影能力,即模型对未见物体类型的泛化能力,发现传统模型如YOLO虽在训练集上达80%准确率,但在未见物体上骤降至15%;(ii)提出Grasp-LLaVA,一种基于视觉语言的抓取模型,通过类人推理根据物体物理特性预测抓取方式,在未见物体类型上实现50.2%的准确率,显著优于当前最优模型的36.7%;最后,为弥合性能与延迟的差距,提出混合抓取网络(HGN),采用边缘-云端协同部署,实现在边缘端快速抓取估计、云端高精度推理作为容错机制,有效拓展延迟-准确率帕累托边界。结合置信度校准(DC)的HGN可动态切换边缘与云端模型,使未见物体的语义投影准确率提升至42.3%(+5.6%),速度提升3.5倍。在真实场景样本混合测试中,平均准确率达86%(较纯边缘方案提升12.2%),且推理速度比单独使用Grasp-LLaVA快2.2倍。
原文摘要 · Abstract (English)
For transradial amputees, robotic prosthetic hands promise to regain the capability to perform daily living activities. To advance next-generation prosthetic hand control design, it is crucial to address current shortcomings in robustness to out of lab artifacts, and generalizability to new environments. Due to the fixed number of object to interact with in existing datasets, contrasted with the virtually infinite variety of objects encountered in the real world, current grasp models perform poorly on unseen objects, negatively affecting users' independence and quality of life. To address this: (i) we define semantic projection, the ability of a model to generalize to unseen object types and show that conventional models like YOLO, despite 80% training accuracy, drop to 15% on unseen objects. (ii) we propose Grasp-LLaVA, a Grasp Vision Language Model enabling human-like reasoning to infer the suitable grasp type estimate based on the object's physical characteristics resulting in a significant 50.2% accuracy over unseen object types compared to 36.7% accuracy of an SOTA grasp estimation model. Lastly, to bridge the performance-latency gap, we propose Hybrid Grasp Network (HGN), an edge-cloud deployment infrastructure enabling fast grasp estimation on edge and accurate cloud inference as a fail-safe, effectively expanding the latency vs. accuracy Pareto. HGN with confidence calibration (DC) enables dynamic switching between edge and cloud models, improving semantic projection accuracy by 5.6% (to 42.3%) with 3.5x speedup over the unseen object types. Over a real-world sample mix, it reaches 86% average accuracy (12.2% gain over edge-only), and 2.2x faster inference than Grasp-LLaVA alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。