构建首个聚焦物体朝向的多模态评测基准,揭示大模型在几何推理上的短板。
Seeing Isn't Orienting: A Cognitively Informed Hierarchical Benchmark for Object Orientation in MLLMs
- 分四维度、粗粒度与细粒度双层设计,隔离朝向与空间干扰因素
- 26个主流模型在细粒度判断上最高仅42.9%,显著低于粗粒度的64.2%
- 适合评估多模态模型在物体朝向理解上的真实几何推理能力
人类对物体朝向的理解是逐步发展的,从识别物体朝向,到跨物体的朝向推理。然而现有视觉语言评测基准大多将朝向与更广泛的时空推理混淆。我们提出认知启发的分层评测基准DORI,将物体朝向作为核心评估目标。DORI将朝向分解为四个维度,每个维度在粗粒度(分类)与细粒度(度量)层面评估,共生成33,656道多选题,基于14个来源的13,652张真实与合成图像。其设计通过物体隔离、标准参考系和结构化提示,有效分离朝向与干扰因素。评估26个先进视觉语言模型发现:在通用空间推理任务表现良好的模型,在物体中心朝向推理上仍接近随机水平。最佳模型在粗粒度判断上达到64.2%,细粒度仅为42.9%,在复合旋转与跨物体参考系转换任务中下降最明显。粗细粒度间显著差距表明模型依赖分类启发式而非几何推理。这些结果确立了物体朝向理解是多模态系统的核心开放挑战。数据集:https://huggingface.co/datasets/appledora/DORI-Benchmark
原文摘要 · Abstract (English)
Humans develop object orientation understanding progressively, from recognizing which way an object faces to reasoning about orientations across multiple objects. Yet existing vision-language benchmarks largely conflate orientation with broader spatial reasoning. We introduce Discriminative Orientation Reasoning Intelligence (DORI), a cognition-informed hierarchical benchmark that establishes object orientation as the primary evaluation target. DORI decomposes orientation into four dimensions, each evaluated at coarse (categorical) and granular (metric) levels, yielding 33,656 multiple-choice questions over 13,652 real-world and synthetic images from 14 sources. Its design isolates orientation from confounding factors through object isolation, standardized reference frames, and structured prompts. Evaluating 26 state-of-the-art vision-language models reveals a consistent limitation: models strong on general spatial benchmarks remain near-random on object-centric orientation reasoning. The best model achieves only 64.2\% on coarse and 42.9\% on granular judgments, with the largest drops on compound rotations and inter-object reference frame shifts. Large coarse-to-granular performance gaps further indicate reliance on categorical heuristics rather than geometric reasoning. These results establish object orientation understanding as a fundamental open challenge for multimodal systems. Dataset: https://huggingface.co/datasets/appledora/DORI-Benchmark
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。