提出分层抽象机制让模型像人一样识别3D物体
Hierarchical Abstraction Enables Human-Like 3D Object Recognition in Deep Learning Models
- 用点云数据训练模型,通过分层抽象捕捉局部与全局几何特征
- 视觉变换器模型在低密度点云下仍保持高准确率,接近人类表现
- 适合关注3D视觉理解与类人认知机制的研究者
人类和深度学习模型都能从稀疏的3D视觉信息(如随机采样的点云)中识别物体。尽管深度学习模型在3D物体识别任务上已达到人类水平,但其是否具备与人类视觉相似的3D形状表征仍不明确。我们假设:在3D形状训练下,模型能形成对局部几何结构的表征,但全局形状表征可能受限。为此,我们系统地设计了两项人类实验,分别操控点密度、物体朝向(实验1)和局部几何结构(实验2)。结果显示,人类在所有条件下均表现稳定。我们对比了基于卷积神经网络(DGCNN)和视觉变换器(Point Transformer)的两类模型,发现点变换器模型比卷积模型更贴近人类表现。这一优势主要源于其支持3D形状分层抽象的机制。
原文摘要 · Abstract (English)
Both humans and deep learning models can recognize objects from 3D shapes depicted with sparse visual information, such as a set of points randomly sampled from the surfaces of 3D objects (termed a point cloud). Although deep learning models achieve human-like performance in recognizing objects from 3D shapes, it remains unclear whether these models develop 3D shape representations similar to those used by human vision for object recognition. We hypothesize that training with 3D shapes enables models to form representations of local geometric structures in 3D shapes. However, their representations of global 3D object shapes may be limited. We conducted two human experiments systematically manipulating point density and object orientation (Experiment 1), and local geometric structure (Experiment 2). Humans consistently performed well across all experimental conditions. We compared two types of deep learning models, one based on a convolutional neural network (DGCNN) and the other on visual transformers (point transformer), with human performance. We found that the point transformer model provided a better account of human performance than the convolution-based model. The advantage mainly results from the mechanism in the point transformer model that supports hierarchical abstraction of 3D shapes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。