用多视角图像训练模型,首次实现人类级3D形状感知。
Human-level 3D shape perception emerges from multi-view learning
- 基于自然场景图像训练神经网络,学习相机位置与视觉深度。
- 模型在3D形状推断任务上达到与人类相同的准确率。
- 无需特定任务训练,可预测人类错误模式和反应时间。
人类能从二维视觉输入中推断物体的三维结构。长期以来,模拟这一能力是视觉智能科学与工程的重要目标,但数十年的计算方法均未达到人类水平。本文提出一种新模型框架,直接从实验刺激预测人类对任意物体的3D形状推断。该框架采用新型神经网络,在自然场景下的多视角图像数据上,通过视觉-空间目标进行训练,学习相机位置、视觉深度等空间信息,无需依赖任何物体相关先验假设。这些信号与人类可用的感官线索相似。我们设计零样本评估方法,测试模型在经典3D感知任务上的表现,并与人类行为对比。结果表明,该模型在无任务特定训练或微调的情况下,首次达到人类水平的3D形状推断准确率。更重要的是,模型输出的独立读出能精准预测人类行为的细微特征,包括错误模式和反应时间,揭示了模型动态与人类感知间的自然对应关系。研究说明,人类级3D感知可从自然化视觉-空间数据的简单、可扩展学习目标中涌现。代码、图像及人类数据详见https://tzler.github.io/human_multiview/
原文摘要 · Abstract (English)
Humans can infer the three-dimensional structure of objects from two-dimensional visual inputs. Modeling this ability has been a longstanding goal for the science and engineering of visual intelligence, yet decades of computational methods have fallen short of human performance. Here we develop a modeling framework that predicts human 3D shape inferences for arbitrary objects, directly from experimental stimuli. We achieve this with a novel class of neural networks trained using a visual-spatial objective over naturalistic sensory data; given a set of images taken from different locations within a natural scene, these models learn to predict spatial information related to these images, such as camera location and visual depth, without relying on any object-related inductive biases. Notably, these visual-spatial signals are analogous to sensory cues readily available to humans. We design a zero-shot evaluation approach to determine the performance of these 'multi-view' models on a well established 3D perception task, then compare model and human behavior. Our modeling framework is the first to match human accuracy on 3D shape inferences, even without task-specific training or fine-tuning. Remarkably, independent readouts of model responses predict fine-grained measures of human behavior, including error patterns and reaction times, revealing a natural correspondence between model dynamics and human perception. Taken together, our findings indicate that human-level 3D perception can emerge from a simple, scalable learning objective over naturalistic visual-spatial data. Code, images, and human data needed to reproduce all analyses can be found at https://tzler.github.io/human_multiview/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。