用神经启发方法直接从事件流中检测6自由度抓取姿势,无需点云重建。
SpikeGrasp: A Benchmark for 6-DoF Grasp Pose Detection from Stereo Spike Streams
- 模仿生物视觉通路,直接处理立体事件流推断抓取姿态
- 在杂乱和无纹理场景中表现优于传统点云方法,数据效率高
- 适合研究类脑机器人控制与低延迟实时抓取系统
大多数机器人抓取系统依赖将传感器数据转换为显式的3D点云,这一计算步骤在生物智能中并不存在。本文探索了一种根本不同的、受神经启发的6-DoF抓取检测范式。我们提出SpikeGrasp框架,模拟生物视觉运动通路,直接处理来自立体事件相机(类似视网膜)的原始异步事件流,以推断抓取姿态。模型融合双目事件流,并使用递归脉冲神经网络,类比高层视觉处理,迭代优化抓取假设,全程不重建点云。为验证该方法,我们构建了一个大规模合成基准数据集。实验表明,SpikeGrasp在杂乱和无纹理场景中超越传统点云基基线,在动态物体抓取中展现出卓越的数据效率。通过证明此端到端神经启发方法的可行性,SpikeGrasp为未来实现自然般流畅高效的操控系统铺平了道路。
原文摘要 · Abstract (English)
Most robotic grasping systems rely on converting sensor data into explicit 3D point clouds, which is a computational step not found in biological intelligence. This paper explores a fundamentally different, neuro-inspired paradigm for 6-DoF grasp detection. We introduce SpikeGrasp, a framework that mimics the biological visuomotor pathway, processing raw, asynchronous events from stereo spike cameras, similarly to retinas, to directly infer grasp poses. Our model fuses these stereo spike streams and uses a recurrent spiking neural network, analogous to high-level visual processing, to iteratively refine grasp hypotheses without ever reconstructing a point cloud. To validate this approach, we built a large-scale synthetic benchmark dataset. Experiments show that SpikeGrasp surpasses traditional point-cloud-based baselines, especially in cluttered and textureless scenes, and demonstrates remarkable data efficiency. By establishing the viability of this end-to-end, neuro-inspired approach, SpikeGrasp paves the way for future systems capable of the fluid and efficient manipulation seen in nature, particularly for dynamic objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。