arXiv:2505.12809cs.LG2025-05被引 3

用动力系统理论建模神经网络内部表示的演化过程

Koopman Autoencoders Learn Neural Representation Dynamics

  • 将神经网络层间表示视为动力系统状态,通过自编码器升维实现线性动态建模
  • 能自然复现神经网络中表示拓扑逐步简化的现象,且在两个数据集上成功实现定向类遗忘
  • 适合研究模型内部表征演化或需可控编辑表示的研究者

本文探讨一个简单问题:能否用动力系统理论建模神经网络的内部变换?我们提出柯普曼自编码器,将神经表示视为动力系统中的状态,学习其从输入到输出的演化规律。该方法通过自编码器对原始状态进行升维,在线性空间中实现动态预测,具有两大优势:首先,升维后可在低维线性空间操作,便于动态编辑;其次,通过正则化自编码目标保留原始表示的拓扑结构。实验表明,这些代理模型能自然复现神经网络中表示拓扑逐步简化的过程。作为实际应用,我们在Yin-Yang和MNIST分类任务中展示了该方法如何实现针对特定类别的定向遗忘。

原文摘要 · Abstract (English)

This paper explores a simple question: can we model the internal transformations of a neural network using dynamical systems theory? We introduce Koopman autoencoders to capture how neural representations evolve through network layers, treating these representations as states in a dynamical system. Our approach learns a surrogate model that predicts how neural representations transform from input to output, with two key advantages. First, by way of lifting the original states via an autoencoder, it operates in a linear space, making editing the dynamics straightforward. Second, it preserves the topologies of the original representations by regularizing the autoencoding objective. We demonstrate that these surrogate models naturally replicate the progressive topological simplification observed in neural networks. As a practical application, we show how our approach enables targeted class unlearning in the Yin-Yang and MNIST classification tasks.

动力系统表示演化类遗忘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。