用稀疏思维生成密集图像,逐步提升超分辨率质量。
Think Sparse, Predict Dense: Continuous Thought Machines for Image Super-Resolution

- 引入连续思维机制,在局部窗口内演化稀疏表示
- 四轮推理后峰值信噪比提升至30.28 dB(Y分量)
- 适合需要精细重建的图像超分辨率任务
连续思维机器在神经元层级的历史与同步表征上引入内部时间维度,随一系列思维周期演进。将该机制扩展至密集视觉预测面临挑战,因图像超分辨率需保持空间证据在每个输出位置可用,而非压缩为单一全局表征。本文提出的窗口级连续思维模型中,思维动态为每个局部窗口生成紧凑摘要表征。DQ-CTM通过结构化低秩、参数高效的紧凑到密集查询转换机制,将此紧凑思维表征转化为窗口对齐的密集查询。窗口内每位置获得独立查询,共享思维动态则逐轮优化密集表征。在超分辨率实例ThinkSR中,编码特征图被划分为无池化的局部视觉窗口,经共享精炼后恢复原始特征域,再解码为高分辨率图像。固定四轮训练期的初步实验显示渐进重建轨迹:T=0时PSNR-Y为28.1045 dB,T=4时达30.2817 dB;PSNR-RGB从26.6271 dB升至28.7781 dB,均方ℓ₁误差由0.034602降至0.023545。全部100张测试图像均从T=1到T=4持续改善。这些结果验证了稀疏潜在思维在密集空间重建中的可行性,并推动更广泛的连续思维架构在密集视觉任务中的应用。
原文摘要 · Abstract (English)
Continuous Thought Machines introduce an internal temporal dimension in which neuron-level histories and synchronization-derived representations evolve over a sequence of thought ticks. Extending this mechanism to dense visual prediction is non-trivial, because tasks such as image super-resolution require spatial evidence to remain available at every output location rather than being compressed into a single global representation. In the proposed window-level use of CTM, the thought dynamics produce a compact summary representation for each local window. DQ-CTM transforms this compact thought representation into window-aligned dense queries through a structured low-rank, parameter-efficient compact-to-dense query mechanism. Each position within a window receives its own query, while shared thought dynamics progressively refine the dense representation across ticks. In its super-resolution instantiation, termed ThinkSR, encoded feature maps are partitioned into local visual windows without token pooling, restored to the original feature field after shared refinement, and decoded into a high-resolution image. Preliminary experiments under a fixed four-tick training horizon reveal a progressive reconstruction trajectory. PSNR-Y increases from 28.1045 dB at $T=0$ to 30.2817 dB at $T=4$, while PSNR-RGB increases from 26.6271 dB to 28.7781 dB and the mean $\ell_1$ error decreases from 0.034602 to 0.023545. All 100 evaluated images improve from $T=1$ to $T=4$. These initial results establish the feasibility of sparse latent thought for dense spatial reconstruction and motivate broader continuous-thought architectures for dense vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。