测试视觉语言模型能否预测图像旋转180度后的内容
Readable Yet Unpredictable: Rotated-Outcome Prediction in Vision-Language Models

- 设计新基准RotOutBench,评估模型对图像旋转结果的预测能力
- 多数模型在直接观察旋转图时能识别内容,但无法从原图推断
- 模型虽能模拟旋转状态,最终输出仍偏向原始文本,存在认知偏差
视觉语言模型能否仅凭原始图像预测出180°平面旋转后的可见内容?我们通过旋转结果预测任务(Rotated-Outcome Prediction)研究这一能力:给定原始图像,模型需回答旋转后将看到或读取的内容,而不能直接观察旋转后的目标。为隔离此能力差距,我们构建了涵盖开放视觉场景与受控图文旋转的配对诊断基准RotOutBench。结果显示:许多视觉语言模型在直接给出原始或旋转图像时可准确识别内容,但仅从原始图像预测旋转结果时表现极差。在受控图文旋转场景中,即使直接阅读准确率很高,预测准确率也几乎降至零。模型级案例分析表明,预测状态可接近旋转图像的读取状态,但最终输出仍倾向于原始字符串。当前视觉语言模型虽能识别被展示的变换视觉状态,却常无法从原视图预测该状态。
原文摘要 · Abstract (English)
Can vision-language models predict what a 180° rotation would reveal from the original image alone? We study this ability through Rotated-Outcome Prediction: given an original image, a model must answer what would be seen or read after a 180° in-plane rotation, without directly observing the rotated target. To isolate this gap, we introduce RotOutBench, a paired diagnostic benchmark spanning open visual cases and controlled text-image rotations. A sharp pattern emerges: many VLMs can recognize the relevant content when directly given either the original or rotated image, yet fail to infer the rotated result from the original image alone. On controlled text-image rotations, predicted-rotation accuracy collapses to near zero even for models with high direct-reading accuracy. A model-level case study further shows that the prediction state can approach a rotated-image reading state, while the final readout still shifts toward the original string. Current VLMs can recognize a transformed visual state when it is shown, but often fail to predict that state from the original view.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。