arXiv:2510.03638cs.LGcs.AI2025-10被引 3

隐式模型通过增加测试时间计算量,逐步提升表达能力与精度。

Expressive Power of Implicit Models: Rich Equilibria and Test-Time Scaling

  • 通过迭代固定参数模块逼近不动点,实现无限深度的紧凑模型。
  • 测试时迭代次数越多,映射复杂度越高,解的质量持续提升并趋于稳定。
  • 适用于图像重建、科学计算等四类任务,尤其适合算力充裕场景。

隐式模型通过迭代单一参数模块至不动点来生成输出,形成一种无限深度且权值共享的网络结构,训练时内存消耗恒定,相比显式模型在相同性能下显著降低内存需求。尽管实证表明这类紧凑模型可通过增加测试时间计算量达到甚至超越大型显式网络的精度,其内在机制仍不清晰。本文通过非参数分析揭示其表达能力,证明一个简单而规则的隐式算子经迭代后可逐步表达更复杂的映射。理论表明,在广泛类别的隐式模型中,表达能力随测试时间计算量增长,最终逼近更丰富的函数类。该理论在图像重建、科学计算、运筹优化和大模型推理四个领域得到验证:随着测试时迭代次数增加,学习映射的复杂度上升,解的质量同步提高并趋于稳定。

原文摘要 · Abstract (English)

Implicit models, an emerging model class, compute outputs by iterating a single parameter block to a fixed point. This architecture realizes an infinite-depth, weight-tied network that trains with constant memory, significantly reducing memory needs for the same level of performance compared to explicit models. While it is empirically known that these compact models can often match or even exceed the accuracy of larger explicit networks by allocating more test-time compute, the underlying mechanism remains poorly understood. We study this gap through a nonparametric analysis of expressive power. We provide a strict mathematical characterization, showing that a simple and regular implicit operator can, through iteration, progressively express more complex mappings. We prove that for a broad class of implicit models, this process lets the model's expressive power scale with test-time compute, ultimately matching a much richer function class. The theory is validated across four domains: image reconstruction, scientific computing, operations research, and LLM reasoning, demonstrating that as test-time iterations increase, the complexity of the learned mapping rises, while the solution quality simultaneously improves and stabilizes.

隐式模型测试时扩展表达能力不动点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。