arXiv:2602.11731cs.CL2026-02

让AI像写代码一样思考视觉问题,提升推理准确性

Thinking with Drafting: Optical Decompression via Logical Reconstruction

  • 用简洁的领域语言作为中间表示,强制模型先写逻辑草稿
  • 在视觉代数任务中准确率显著优于传统方法,避免生成幻觉
  • 适合需要精确推理的科研、数学或工业级视觉分析场景

现有多模态大模型在视觉感知和生成上表现优异,但在复杂推理任务中仍存在精度悖论:视觉系统仅转录符号而忽略逻辑结构,像素生成模型则产生缺乏数学精确性的视觉伪影。为弥合这一鸿沟,我们提出将视觉推理重新定义为光学解压缩——从压缩的视觉标记中重建潜在逻辑结构。基于‘解析即推理’的公理,我们引入思维草稿(TwD)框架,采用极简领域特定语言(DSL)作为基础中间表示。与直接幻觉式回答不同,TwD 强制模型将思维过程转化为可执行代码,实现确定性视觉验证。我们构建了 VisAlg 视觉代数基准进行验证。实验表明,TwD 作为更优的认知支架,建立了一个闭环系统:视觉生成不再只是创造性输出,而是逻辑验证工具,为视觉推理提供通用化路径。

原文摘要 · Abstract (English)

Existing multimodal large language models have achieved high-fidelity visual perception and exploratory visual generation. However, a precision paradox persists in complex reasoning tasks: optical perception systems transcribe symbols without capturing logical topology, while pixel-based generative models produce visual artifacts lacking mathematical exactness. To bridge this gap, we propose that reasoning over visual inputs be reconceptualized as optical decompression-the process of reconstructing latent logical structures from compressed visual tokens. Guided by the axiom that Parsing is Reasoning, we introduce Thinking with Drafting (TwD), which utilizes a minimalist Domain-Specific Language (DSL) as a grounding intermediate representation. Unlike standard approaches that hallucinate answers directly, TwD forces the model to draft its mental model into executable code, rendering deterministic visual proofs for self-verification. To validate this, we present VisAlg, a visual algebra benchmark. Experiments demonstrate that TwD serve as a superior cognitive scaffold. Our work establishes a closed-loop system where visual generation acts not as a creative output but as a logical verifier, offering a generalizable path for visual reasoning.

视觉推理逻辑验证代码生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。