实时动态抓取新方法,59毫秒延迟仍能精准响应用户指令。
SPGrasp: Spatiotemporal Prompt-driven Grasp Synthesis in Dynamic Scenes
- 融合时空提示与上下文,实现端到端快速抓取推断。
- 在多个数据集上抓取准确率超90%,帧延迟仅73.1毫秒。
- 适合需要低延迟交互的机器人抓取场景,如动态物体操作。
实时交互式动态物体抓取仍具挑战,因现有方法难以兼顾低延迟推理与提示响应能力。为此,我们提出SPGrasp(时空提示驱动的动态抓取合成),基于Segment Anything Model v2(SAMv2)扩展用于视频流抓取估计。核心创新在于将用户提示与时空上下文结合,实现端到端延迟低至59毫秒的同时保证动态物体的时间一致性。基准测试中,SPGrasp在OCID数据集上达到90.6%实例级抓取准确率,在Jacquard上达93.8%。在挑战性GraspNet-1Billion数据集连续追踪任务中,每帧延迟73.1毫秒,准确率92.0%,相较前序可提示方法RoG-SAM降低58.5%延迟,且保持竞争力。真实世界实验涉及13个移动物体,交互抓取成功率达94.8%。结果表明,SPGrasp有效解决了动态抓取中延迟与交互性的权衡问题。
原文摘要 · Abstract (English)
Real-time interactive grasp synthesis for dynamic objects remains challenging as existing methods fail to achieve low-latency inference while maintaining promptability. To bridge this gap, we propose SPGrasp (spatiotemporal prompt-driven dynamic grasp synthesis), a novel framework extending segment anything model v2 (SAMv2) for video stream grasp estimation. Our core innovation integrates user prompts with spatiotemporal context, enabling real-time interaction with end-to-end latency as low as 59 ms while ensuring temporal consistency for dynamic objects. In benchmark evaluations, SPGrasp achieves instance-level grasp accuracies of 90.6% on OCID and 93.8% on Jacquard. On the challenging GraspNet-1Billion dataset under continuous tracking, SPGrasp achieves 92.0% accuracy with 73.1 ms per-frame latency, representing a 58.5% reduction compared to the prior state-of-the-art promptable method RoG-SAM while maintaining competitive accuracy. Real-world experiments involving 13 moving objects demonstrate a 94.8% success rate in interactive grasping scenarios. These results confirm SPGrasp effectively resolves the latency-interactivity trade-off in dynamic grasp synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。