arXiv:2502.09297cs.LG2025-02ICML被引 7

神经网络在多任务下能学习到数据生成的潜在变量,但依赖模型结构。

When Do Neural Networks Learn World Models?

  • 通过低阶偏差模型,在多任务设置中可还原潜在生成变量。
  • 即使代理任务是潜在变量的复杂非线性函数,仍能准确恢复。
  • 适用于研究自监督学习与分布外泛化等问题的学者。

人类能够构建捕捉数据生成过程的世界模型。神经网络是否也能学习类似的模型仍是开放问题。本文首次提供了该问题的理论结果:在多任务设置中,具有低度偏差的模型在温和假设下可准确恢复潜在数据生成变量,即使代理任务涉及潜在变量的复杂非线性函数。然而,这种恢复对模型架构敏感。分析基于布尔任务解的傅里叶-沃尔什变换,引入了分析可逆布尔变换的新技术,可能具有独立研究价值。我们展示了这些结果的算法启示,并将其与自监督学习、分布外泛化及大语言模型中的线性表示假设等研究方向相联系。

原文摘要 · Abstract (English)

Humans develop world models that capture the underlying generation process of data. Whether neural networks can learn similar world models remains an open problem. In this work, we present the first theoretical results for this problem, showing that in a multi-task setting, models with a low-degree bias provably recover latent data-generating variables under mild assumptions--even if proxy tasks involve complex, non-linear functions of the latents. However, such recovery is sensitive to model architecture. Our analysis leverages Boolean models of task solutions via the Fourier-Walsh transform and introduces new techniques for analyzing invertible Boolean transforms, which may be of independent interest. We illustrate the algorithmic implications of our results and connect them to related research areas, including self-supervised learning, out-of-distribution generalization, and the linear representation hypothesis in large language models.

世界模型多任务学习理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。