arXiv:2508.13968cs.CVcs.AI2025-08Conference of the …被引 6

测试大模型识图旋转能力,发现多数模型难分90°和270°。

RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation

  • 构建350张图像的旋转识别基准RotBench,含生活、人物、风景类
  • 主流模型能认0°正向,180°倒向,但90°与270°无法可靠区分
  • 提示词优化和微调对90°/270°识别帮助有限,适合空间推理研究者

我们探究多模态大语言模型(MLLMs)在识别输入图像0°、90°、180°、270°旋转方面的能力。该任务需强视觉推理能力以捕捉旋转线索并理解图像中空间关系。为此,我们引入RotBench,一个由350张人工筛选的图像构成的基准,涵盖生活、人物和风景类图像。尽管任务看似简单,但多个最先进的开源及专有MLLM(包括GPT-5、o3和Gemini-2.5-Pro)未能可靠识别图像旋转。提供辅助信息(如描述、深度图)或使用思维链提示仅带来微弱且不一致的改进。结果表明,大多数模型可可靠识别正向(0°)图像,某些模型可识别倒置(180°)图像,但无一能可靠区分90°与270°旋转。同时展示不同方向图像可使推理模型表现小幅提升,而投票机制则改善了弱模型性能。进一步实验显示,微调虽显著提升180°识别准确率,却未提升90°/270°分辨能力。整体揭示了当前MLLM空间推理能力与人类感知间的显著差距。

原文摘要 · Abstract (English)

We investigate to what extent Multimodal Large Language Models (MLLMs) can accurately identify the orientation of input images rotated 0°, 90°, 180°, and 270°. This task demands robust visual reasoning capabilities to detect rotational cues and contextualize spatial relationships within images, regardless of their orientation. To evaluate MLLMs on these abilities, we introduce RotBench, a 350-image manually-filtered benchmark comprising lifestyle, portrait, and landscape images. Despite the relatively simple nature of this task, we show that several state-of-the-art open and proprietary MLLMs, including GPT-5, o3, and Gemini-2.5-Pro, do not reliably identify rotation in input images. Providing models with auxiliary information -- including captions, depth maps, and more -- or using chain-of-thought prompting offers only small and inconsistent improvements. Our results indicate that most models are able to reliably identify right-side-up (0°) images, while certain models are able to identify upside-down (180°) images. None can reliably distinguish between 90° and 270° rotated images. Simultaneously showing the image rotated in different orientations leads to moderate performance gains for reasoning models, while a modified setup using voting improves the performance of weaker models. We further show that fine-tuning does not improve models' ability to distinguish 90° and 270° rotations, despite substantially improving the identification of 180° images. Together, these results reveal a significant gap between MLLMs' spatial reasoning capabilities and human perception in identifying rotation.

多模态空间推理图像旋转评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。