arXiv:2606.09585cs.AI2026-06

用图像代替文字做推理,更高效且统一。

Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text

论文配图:Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text
图 1 · 摘自论文原文
  • 让图像独立承载推理过程,无需文字辅助。
  • 在语言任务上减少28.57%推理令牌,多模态任务减16%。
  • 适合追求高效推理与视觉化表达的研究者。

链式思维(CoT)提升了大语言模型(LLMs)性能,并被拓展至多模态大语言模型(MLLMs)。近期研究进一步从文本驱动的多模态推理转向跨模态交织推理,允许中间步骤融合文本解释与视觉证据。本文提出更具突破性的观点:图像能否独立作为语言与多模态任务的推理介质?为此,我们引入光学推理(Optical Reasoning),将图像视为独立推理载体。通过两种实现方式:基于排版的光学推理优化视觉布局以紧凑呈现推理过程;基于图形的光学推理将文字与图形元素组合成结构化视觉推理。在数学、科学及跨模态推理基准测试中,光学推理表现可媲美甚至超越传统文本推理,在语言任务上平均减少28.57%推理令牌,多模态任务减少16%,达到文本推理1.96倍的令牌效率。结果表明,图像能有效且高效地编码推理过程,并提供统一的视觉推理画布。

原文摘要 · Abstract (English)

Chain-of-Thought (CoT) improves the performance of Large Language Models (LLMs) and has been extended to Multimodal Large Language Models (MLLMs). More recent work further moves from text-based multimodal reasoning toward interleaved-modal reasoning, where intermediate steps can incorporate both textual rationales and visual evidence. In this work, we propose a bolder and more ambitious idea: could images alone serve as the reasoning medium for both language and multimodal tasks? To explore this, we propose optical reasoning, which treats images as a standalone reasoning medium. We instantiate this concept with two variants: typographic-based optical reasoning, which optimizes visual layouts for compact rationale rendering, and graphical-based optical reasoning, which composes text and graphical elements into structured visual rationales. Across mathematical, scientific, and interleaved-modal reasoning benchmarks, optical reasoning can match or even exceed traditional text reasoning while reducing reasoning tokens by an average of 28.57% on language tasks and 16% on multimodal tasks, achieving 1.96 times the token efficiency of text reasoning. These results show that images can effectively and efficiently encode rationales while providing a unified visual canvas for reasoning.

视觉推理图像编码令牌效率多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。