arXiv:2604.25065cs.CV2026-04

提出新基准ShapeY,评估模型对形状的识别能力

ShapeY: A Principled Framework for Measuring Shape Recognition Capacity via Nearest-Neighbor Matching

论文配图:ShapeY: A Principled Framework for Measuring Shape Recognition Capacity via Nearest-Neighbor Matching
图 1 · 摘自论文原文
  • 用最近邻匹配测试模型在不同视角下的形状区分能力
  • 321个模型中顶尖模型仍难在视角和外观变化下稳定识别
  • 适合研究视觉模型泛化性与形状不变性的研究人员

人类物体识别高度依赖形状线索及跨三维视角的识别能力。而深度网络常依赖纹理、背景等非形状线索,导致泛化与鲁棒性不足。为此,我们提出ShapeY——一种全新的基准框架,用于评估物体识别系统基于形状的能力。ShapeY包含68,200张200个3D物体从多个视角生成的灰度图像,并可施加非形状外观变化。通过最近邻匹配任务,该框架聚焦于探测模型嵌入空间的细粒度结构,评估不同视角和外观变化下,物体视图是否按3D形状相似性聚类。提供包括错误率曲线、视角调优曲线、正负匹配得分直方图及最佳匹配排列网格在内的定量与定性分析结果,全面评估模型的形状理解能力。对321个预训练网络的测试显示,即便先进模型在视角与外观变化下也难以保持一致性能,且存在明显形状差异的物体仍会误匹配。ShapeY为推动人工视觉系统向类人形状识别迈进提供了原则性框架,强调解耦与不变编码的重要性。

原文摘要 · Abstract (English)

Object recognition (OR) in humans relies heavily on shape cues and the ability to recognize objects across varying 3D viewpoints. Unlike humans, deep networks often rely on non-shape cues such as texture and background, leading to vulnerabilities in generalization and robustness. To address this gap, we introduce ShapeY, a novel and principled benchmarking framework designed to evaluate shape-based recognition capability in OR systems. ShapeY comprises 68,200 grayscale images of 200 3D objects rendered from multiple viewpoints and optionally subjected to non-shape ``appearance'' changes. Using a nearest-neighbor matching task, ShapeY specifically probes the fine-grained structure of an OR system's embedding space by evaluating whether object views are clustered by 3D shape similarity across varying 3D viewpoints and other non-shape changes. ShapeY provides a suite of quantitative and qualitative performance readouts, including error rate graphs, viewpoint tuning curves, histograms of positive and negative matching scores, and grids showing ordered best matches, which together offer a comprehensive evaluation of an OR system's shape understanding capability. Testing of 321 pre-trained networks with diverse architectures reveals significant challenges in achieving robust shape-based recognition: even state-of-the-art models struggle to generalize consistently across 3D viewpoint and appearance changes, and are prone to infrequent but egregious matches of objects of obviously completely different shape. ShapeY establishes a principled framework for advancing artificial vision systems toward human-like shape recognition capabilities, emphasizing the importance of disentangled and invariant object encodings.

形状识别基准测试视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。