arXiv:2504.19256cs.CVcs.RO2025-04被引 1

轻量级多视角融合模型提升机器人3D物体识别准确率

LM-MCVT: A Lightweight Multi-modal Multi-view Convolutional-Vision Transformer Approach for 3D Object Recognition

  • 结合卷积与视觉变压器,分层提取多视角特征
  • 在ModelNet40上达95.6%准确率,优于现有方法
  • 适用于真实场景与合成数据,鲁棒性强

在餐厅、家庭和仓库等以人为中心的环境中,机器人常面临3D物体识别困难的问题,这源于环境复杂性和物体形状多样性。本文提出一种轻量级多模态多视角卷积-视觉变换器网络(LM-MCVT),以提升机器人在3D物体识别中的表现。该方法采用基于全局熵的嵌入融合(GEEF)策略,高效整合多视角信息。LM-MCVT架构包含预处理和中层卷积编码器,以及局部与全局视觉变换器,有效增强特征提取与识别精度。我们在合成的ModelNet40数据集上评估,四视角设置下达到95.6%的识别准确率,超越现有最优方法。为进一步验证有效性,我们在真实世界数据集OmniObject3D上采用相同配置进行五折交叉验证,结果持续表现出色,证明该方法在合成与真实3D数据上的鲁棒性。

原文摘要 · Abstract (English)

In human-centered environments such as restaurants, homes, and warehouses, robots often face challenges in accurately recognizing 3D objects. These challenges stem from the complexity and variability of these environments, including diverse object shapes. In this paper, we propose a novel Lightweight Multi-modal Multi-view Convolutional-Vision Transformer network (LM-MCVT) to enhance 3D object recognition in robotic applications. Our approach leverages the Globally Entropy-based Embeddings Fusion (GEEF) method to integrate multi-views efficiently. The LM-MCVT architecture incorporates pre- and mid-level convolutional encoders and local and global transformers to enhance feature extraction and recognition accuracy. We evaluate our method on the synthetic ModelNet40 dataset and achieve a recognition accuracy of 95.6% using a four-view setup, surpassing existing state-of-the-art methods. To further validate its effectiveness, we conduct 5-fold cross-validation on the real-world OmniObject3D dataset using the same configuration. Results consistently show superior performance, demonstrating the method's robustness in 3D object recognition across synthetic and real-world 3D data.

3D识别多视图轻量化机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。