用双流网络融合视觉触觉信息,提升机器人对物体形状硬度的识别准确率。
Two-stream network-driven vision-based tactile sensor for object feature extraction and fusion perception
- 双流设计分别提取物体内外特征,融合深度与接触力数据。
- 硬度识别准确率达98.0%,实际抓取场景下整体识别超98.5%。
- 适合需要精细感知的机器人抓取与具身智能任务。
触觉感知对具身智能机器人识别物体至关重要。基于视觉的触觉传感器通过高空间分辨率多维提取物体物理属性,但产生大量冗余信息;单一维度提取缺乏有效融合,难以全面表征物体特性,制约识别精度提升。为此,本文提出一种双流网络驱动的特征提取与融合感知策略,采用分布式方法分别提取物体内部与外部特征:通过三维重建获取深度图信息,同时利用接触力数据测量硬度信息。经卷积神经网络(CNN)提取特征后,应用加权融合生成更丰富有效的特征表示。在不同形状与硬度物体的标准测试中,力预测误差为0.06 N(量程12 N),硬度识别准确率达98.0%,形状识别准确率达93.75%。融合算法使实际抓取场景下的物体识别准确率超过98.5%。该方法聚焦于物体物理属性感知,增强人工触觉系统从感知到认知的跃迁能力,适用于具身感知应用场景。
原文摘要 · Abstract (English)
Tactile perception is crucial for embodied intelligent robots to recognize objects. Vision-based tactile sensors extract object physical attributes multidimensionally using high spatial resolution; however, this process generates abundant redundant information. Furthermore, single-dimensional extraction, lacking effective fusion, fails to fully characterize object attributes. These challenges hinder the improvement of recognition accuracy. To address this issue, this study introduces a two-stream network feature extraction and fusion perception strategy for vision-based tactile systems. This strategy employs a distributed approach to extract internal and external object features. It obtains depth map information through three-dimensional reconstruction while simultaneously acquiring hardness information by measuring contact force data. After extracting features with a convolutional neural network (CNN), weighted fusion is applied to create a more informative and effective feature representation. In standard tests on objects of varying shapes and hardness, the force prediction error is 0.06 N (within a 12 N range). Hardness recognition accuracy reaches 98.0%, and shape recognition accuracy reaches 93.75%. With fusion algorithms, object recognition accuracy in actual grasping scenarios exceeds 98.5%. Focused on object physical attributes perception, this method enhances the artificial tactile system ability to transition from perception to cognition, enabling its use in embodied perception applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。