arXiv:2602.05570cs.AI2026-02

让视觉语言模型像人一样通过试错迭代改进几何推理能力

TangramSR: Can Vision-Language Models Reason in Continuous Geometric Space?

  • 用上下文学习+奖励反馈循环模拟人类试错过程
  • 无训练情况下,中等三角形任务的重合率从0.63提升至0.932
  • 为自进化AI在连续空间推理中提供可落地的解决方案

人类在解决拼图类空间推理任务时,依赖心理旋转、迭代修正和视觉反馈等认知机制。受此启发,本文设计了一种模拟人类认知过程的框架。然而,对五种代表性视觉语言模型(VLMs)的全面实验显示,其在连续几何推理上存在系统性缺陷:单块任务平均交并比(IoU)仅0.41,两块组合任务降至0.23,远低于人类儿童水平。本文解决自进化AI的核心挑战——能否在不更新参数的情况下,在测试时迭代优化预测?提出一种无需训练的验证-修正代理框架,结合上下文学习(ICL)与奖励引导反馈环,通过递归修正循环基于几何一致性反馈持续优化预测。在中等三角形案例中,IoU从0.63提升至0.932,证明融入人类启发的迭代修正机制,可通过ICL与奖励循环显著增强VLM的几何推理能力,使自进化AI从愿景走向实践。

原文摘要 · Abstract (English)

Humans excel at spatial reasoning tasks like Tangram puzzle assembly through cognitive processes involving mental rotation, iterative refinement, and visual feedback. Inspired by how humans solve Tangram puzzles through trial-and-error, observation, and correction, we design a framework that models these human cognitive mechanisms. However, comprehensive experiments across five representative Vision-Language Models (VLMs) reveal systematic failures in continuous geometric reasoning: average IoU of only 0.41 on single-piece tasks, dropping to 0.23 on two-piece composition, far below human performance where children can complete Tangram tasks successfully. This paper addresses a fundamental challenge in self-improving AI: can models iteratively refine their predictions at test time without parameter updates? We introduce a test-time self-refinement framework that combines in-context learning (ICL) with reward-guided feedback loops, inspired by human cognitive processes. Our training-free verifier-refiner agent applies recursive refinement loops that iteratively self-refine predictions based on geometric consistency feedback, achieving IoU improvements from 0.63 to 0.932 on medium-triangle cases without any model retraining. This demonstrates that incorporating human-inspired iterative refinement mechanisms through ICL and reward loops can substantially enhance geometric reasoning in VLMs, moving self-improving AI from promise to practice in continuous spatial domains. Our work is available at this anonymous link https://anonymous.4open.science/r/TangramVLM-F582/.

视觉语言模型几何推理自进化AI上下文学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。