arXiv:2602.07061cs.LGcs.AI2026-02被引 1

TACIT通过像素空间扩散模型实现可视化的视觉推理,揭示了类人类的顿悟式思维过程。

TACIT: Transformation-Aware Capturing of Implicit Thought

  • 在像素空间使用修正流建模,直接可视化每一步推理过程
  • 100万合成迷宫对上训练损失降192倍,重建误差降低22.7倍
  • 发现推理存在68%无可见解的“潜伏期”,后在2%时间内突然整体显现

我们提出TACIT(Transformation-Aware Capturing of Implicit Thought),一种基于扩散模型的可解释视觉推理变压器。与依赖语言的推理系统不同,TACIT完全在像素空间中运行,采用修正流(rectified flow),可在每次推理步骤中直接可视化推理过程。我们在迷宫求解任务上验证该方法,模型学习将未解迷宫图像转化为解谜结果。在100万组合成迷宫对上的关键结果包括:训练损失在100个周期内降低192倍,与真实解的L2距离改善22.7倍,仅需10次欧拉步数(对比典型扩散模型的100–1000步)。定量分析揭示显著的相变现象:解在68%的变换过程中不可见(召回率为零),随后在t=0.70处仅用2%的流程即突然整体出现。最令人惊讶的是,所有样本均在同一时刻、全空间同步显现,排除了逐步路径构建的可能性,表明其为整体性而非算法式推理。这种“顿悟时刻”模式——长期潜伏后突然涌现——与人类认知中的洞察现象高度相似。像素空间设计结合无噪声流匹配,为理解神经网络如何发展出语言之前、隐含于底层的推理策略提供了基础。

原文摘要 · Abstract (English)

We present TACIT (Transformation-Aware Capturing of Implicit Thought), a diffusion-based transformer for interpretable visual reasoning. Unlike language-based reasoning systems, TACIT operates entirely in pixel space using rectified flow, enabling direct visualization of the reasoning process at each inference step. We demonstrate the approach on maze-solving, where the model learns to transform images of unsolved mazes into solutions. Key results on 1 million synthetic maze pairs include: - 192x reduction in training loss over 100 epochs - 22.7x improvement in L2 distance to ground truth - Only 10 Euler steps required (vs. 100-1000 for typical diffusion models) Quantitative analysis reveals a striking phase transition phenomenon: the solution remains invisible for 68% of the transformation (zero recall), then emerges abruptly at t=0.70 within just 2% of the process. Most remarkably, 100% of samples exhibit simultaneous emergence across all spatial regions, ruling out sequential path construction and providing evidence for holistic rather than algorithmic reasoning. This "eureka moment" pattern -- long incubation followed by sudden crystallization -- parallels insight phenomena in human cognition. The pixel-space design with noise-free flow matching provides a foundation for understanding how neural networks develop implicit reasoning strategies that operate below and before language.

视觉推理扩散模型可解释性顿悟机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。