测试视觉语言模型在相机倾斜和物体遮挡下的几何推理可靠性
Tri-Bench: Stress-Testing VLM Reliability on Spatial Reasoning under Camera Tilt and Object Interference
- 设计三角形任务基准,分离几何推理与干扰因素
- 平均准确率仅69%,相机倾斜使性能下降4.1%
- 对特殊三角形类型识别失败,依赖图像平面而非提示引导
可验证的几何推理是可信可控智能体AI的关键。尽管表现强劲,视觉语言模型(VLMs)在真实场景变化下常失效。我们提出Tri-Bench,一个聚焦平面三角形问题的紧凑基准,隔离相对几何推理,并施加两个部署关键因素:相机位姿(平面与倾斜)和物体干扰(10种日常物品)。为测试可验证性与控制能力,使用单一固定提示,其显式描述周围方形边界,可通过单应性正确求解。评估六项简单任务,目标为二元或连续值,发现相对于3D真值的平均准确率仅为~69%(最佳~75%,最差~64%)。相同回答在图像平面2D投影中更接近,平均准确率~72%。所有四款VLM均一致失败,对少数形状类别(等边、等腰、直角三角形)识别准确率降至~0%。此外,相机倾斜导致整体准确率下降~4.1%,表明模型未能利用提示中的参考框架提示,仍依赖2D图像线索。最后发现,物体干扰对准确率无显著影响。
原文摘要 · Abstract (English)
Verifiable geometric reasoning is a critical component for trustworthy and controllable agentic AI. Despite impressive capabilities, Vision-Language Models (VLMs) often fail under realistic scene changes. We present Tri-Bench, a compact benchmark of planar triangle problems that isolates relative geometric reasoning while stressing two deployment-critical factors: camera pose (planar vs. tilted) and scene context via object interference (10 everyday objects). To test verifiability and control, we evaluate four recent VLMs using a single, fixed prompt whose guardrail explicitly describes a surrounding square border, enabling correct answers via homography. We evaluate six simple tasks over binary and continuous targets, and observe that the overall accuracy with respect to 3D ground truth is modest, ~69% on average (best ~75%, worst ~64%). The same responses align even more closely with 2D projections in the image plane, where mean accuracy is ~72%. All four VLMs consistently fail, with accuracy falling to ~0%, on recognizing minority shape classes (equilateral, isosceles, right-angled triangles). Additionally, overall VLM accuracy degrades by ~4.1% under camera tilt. This demonstrates that models fail to correctly utilize the explicit frame-of-reference hint provided in the prompt and default to 2D image plane cues. Finally, we find that object interference has no significant effect on VLM accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。