arXiv:2412.20631cs.CV2024-12被引 26

让AI像人一样一步步看图,提升几何图形理解能力。

Slow Perception: Let's Perceive Geometric Figures Step-by-step

  • 分步解析图形:将复杂几何图拆解为点线单元,逐步重建结构。
  • 使用感知尺追踪每条线段,避免跳步错误,准确率显著提升。
  • 越慢推理越好,符合人类逐步观察的思维规律,适合几何任务。

近期,'视觉o1'设计引发关注,期望其通过缓慢思考解决视觉推理任务,尤其是几何数学问题。然而,当前大视觉语言模型(LVLM)连准确复制几何图形都难以做到,更不用说理解其中复杂的内在逻辑与空间关系。我们认为,准确复制(强感知)是实现视觉o1的第一步。为此,我们提出‘慢感知’(SP)概念,引导模型如人类般逐步感知基本点线组合,渐进重构复杂几何结构。SP包含两阶段:一是感知分解,将复杂图形拆解为基本简单单元,统一几何表示;二是感知流,认识到精确追踪线条并非易事,通过提出的‘感知尺’逐笔追踪线段,避免‘长视觉跳跃’。令人惊讶的是,这种类人感知方式遵循推理时间缩放定律——越慢越好。过去研究致力于加速模型感知,而我们反其道而行之,使模型逐步、细致地读图。

原文摘要 · Abstract (English)

Recently, "visual o1" began to enter people's vision, with expectations that this slow-thinking design can solve visual reasoning tasks, especially geometric math problems. However, the reality is that current LVLMs (Large Vision Language Models) can hardly even accurately copy a geometric figure, let alone truly understand the complex inherent logic and spatial relationships within geometric shapes. We believe accurate copying (strong perception) is the first step to visual o1. Accordingly, we introduce the concept of "slow perception" (SP), which guides the model to gradually perceive basic point-line combinations, as our humans, reconstruct complex geometric structures progressively. There are two-fold stages in SP: a) perception decomposition. Perception is not instantaneous. In this stage, complex geometric figures are broken down into basic simple units to unify geometry representation. b) perception flow, which acknowledges that accurately tracing a line is not an easy task. This stage aims to avoid "long visual jumps" in regressing line segments by using a proposed "perceptual ruler" to trace each line stroke-by-stroke. Surprisingly, such a human-like perception manner enjoys an inference time scaling law -- the slower, the better. Researchers strive to speed up the model's perception in the past, but we slow it down again, allowing the model to read the image step-by-step and carefully.

视觉推理几何理解慢思考感知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。