发现视觉推理模型依赖坐标网格漏洞,换极坐标后性能暴跌。
The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space

- 将53个视觉推理任务转为极坐标空间,打破模型依赖的正交结构
- 14个主流模型在极坐标上准确率从70%-83%降至31%-39%
- 揭示模型缺乏拓扑不变的视觉推理能力,适合研究多模态泛化者
当前多模态大语言模型在标准视觉推理基准上表现接近饱和,但其高分是否反映真实视觉理解仍存疑问。我们识别出一种普遍存在的漏洞——‘笛卡尔捷径’:现有基准普遍采用正交网格布局,可被轻易离散化为显式文本坐标,导致模型系统性地利用文本推理辅助视觉解题。为系统性消除此捷径,我们提出Polaris-Bench,将53个视觉推理任务重构至极坐标空间,并保留对应的笛卡尔版本作为参照,同时保持一致的逻辑约束与任务语义——从根本上打破模型所依赖的正交先验。对14个前沿多模态大模型的全面评估显示,这些在笛卡尔布局上取得70%–83%准确率的模型,在极坐标等价任务中下降至31%–39%,且即使在逻辑完全等价条件下,性能衰退依然持续。此外,笛卡尔布局上的推理优势在极坐标中显著减弱。这些发现暴露了当前多模态大模型在拓扑不变视觉推理方面的关键缺陷。
原文摘要 · Abstract (English)
As current Multimodal Large Language Models rapidly saturate canonical visual reasoning benchmarks, a key question emerges: do these strong scores genuinely reflect robust visual understanding? We identify a pervasive vulnerability, the Cartesian Shortcut: visual reasoning benchmarks prevalently build on orthogonal grid-based layouts that can be readily discretized into explicit textual coordinates. Models systematically exploit this property, heavily leveraging text-based deductive reasoning to assist visual problem-solving. To systematically dismantle this shortcut, we introduce Polaris-Bench, which re-formulates 53 visual reasoning tasks in Polar coordinate space with paired Cartesian counterparts as reference, while preserving consistent logical constraints and task semantics -- thus fundamentally breaking the orthogonal prior that models exploit. Comprehensive evaluation across $14$ state-of-the-art MLLMs reveals that frontier models achieving $70$--$83\%$ on Cartesian layouts collapse to $31$--$39\%$ on Polar equivalents, with degradation persisting even under complete logical equivalence. Moreover, reasoning gains observed on Cartesian layouts are severely diminished on Polar equivalents. These findings expose a critical deficiency in current MLLMs: the lack of topology-invariant visual reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。