arXiv:2603.11410cs.CV2026-03被引 1

新基准揭示大模型在物体朝向理解上的系统性失败。

Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs Supplementary

  • 构建分层朝向推理基准DORI,分离朝向与位置等干扰因素。
  • 24个主流模型在朝向任务上最高仅54.2%准确率,且对复合旋转表现差。
  • 适合关注机器人操作、3D重建的开发者和研究者。

人类学习物体朝向是逐步的:从识别方向,到心理旋转,再到物体间朝向关系推理。现有视觉语言基准大多将朝向与位置及场景理解混为一谈。本文提出认知基础的层级化基准DORI,以物体朝向为核心目标。受人类朝向认知阶段启发,DORI将朝向分解为四个维度,在粗粒度(分类)与细粒度(度量)两个层次评估。基准包含13,652张图像,来自14个来源,涵盖67种物体类别,生成33,656道多选题,覆盖真实世界与合成场景。通过边界框隔离、标准化空间参考系和结构化提示,有效排除物体识别难度、场景杂乱和语言模糊等干扰。评估24个前沿视觉语言模型发现:在通用空间基准表现良好的模型,在面向物体的朝向任务上近乎随机。最佳模型在粗粒度任务上仅达54.2%,细粒度为45.0%,尤其在复合旋转和物体间参考系变换中表现最差。巨大的粗细差距表明模型依赖分类启发式而非几何推理,这一缺陷被现有基准掩盖。结果表明,朝向理解仍是多模态系统的未解难题,对机器人操控、3D场景重建和人机交互具有深远影响。

原文摘要 · Abstract (English)

Humans learn object orientation progressively, from recognizing which way an object faces, to mentally rotating it, to reasoning about orientations between objects. Current vision-language benchmarks largely conflate orientation with position and general scene understanding. We introduce Discriminative Orientation Reasoning Intelligence (DORI), a cognitively grounded hierarchical benchmark that makes object orientation the primary target. Inspired by stages of human orientation cognition, DORI decomposes orientation into four dimensions, each evaluated at coarse (categorical) and granular (metric) levels. Composed from 13,652 images across 14 sources, DORI provides 33,656 multiple-choice questions covering 67 object categories in real-world and synthetic settings. Its coarse-to-granular design isolates orientation from confounds such as object recognition difficulty, scene clutter, and linguistic ambiguity via bounding-box isolation, standardized spatial reference frames, and structured prompts. Evaluating 24 state-of-the-art vision-language models shows a clear pattern: models that perform well on general spatial benchmarks are near-random on object-centric orientation tasks. The best models reach only 54.2% on coarse and 45.0% on granular judgments, with largest failures on compound rotations and shifts in inter-object reference frames. Large coarse-to-granular gaps reveal reliance on categorical heuristics rather than geometric reasoning, a limitation hidden by existing benchmarks. These results identify orientation understanding as an unsolved challenge for multimodal systems, with implications for robotic manipulation, 3D scene reconstruction, and human-AI interaction.

多模态朝向理解基准测试机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。