arXiv:2609.03495cs.LGstat.ML2026-09

训练好的自编码器参数可作数据样本的向量表示。

Spectral characteristics of autoencoder parameters as a vector representation of data

论文配图:Spectral characteristics of autoencoder parameters as a vector representation of data
图 1 · 摘自论文原文
  • 用参数矩阵的谱特征构建数据向量表示
  • 在CIFAR-10和FashionMNIST上区分效果优异
  • 无需原始样本,适合模型指纹分析

本文研究自编码器模型参数与训练数据统计特性之间的关系。自编码器采用编码器-解码器架构,通过压缩隐状态重建输入数据。本文提出,模型参数可视为对应样本的稠密向量表示。为验证该假设,开展理论与实验研究,基于自编码器参数矩阵的谱特征构建向量表示。理论分析表明,模型参数矩阵的奇异值与训练数据协方差矩阵的特征值相关,确保数据空间与参数空间间的信息传递。在CIFAR-10和FashionMNIST数据集上的实验结果表明,所生成的向量表示可在不依赖复杂向量生成算法或原始样本的情况下,高精度区分在不同数据子集上训练的模型。结果表明,训练后的自编码器参数可作为样本的有效表示。

原文摘要 · Abstract (English)

This paper examines the relationship between the parameters of autoencoder models and the statistical properties of the data on which they are trained. Autoencoders are defined as models with an encoder-decoder architecture, trained to reconstruct input data through a compressed latent representation. It is proposed that the model parameters can be viewed as a dense vector representation of the corresponding sample. To test this hypothesis, a theoretical and experimental study is conducted in which a vector representation is formed based on the spectral characteristics of the autoencoder parameter matrices. Theoretical analysis shows that the singular values of the model parameter matrices are related to the eigenvalues of the covariance matrix of the training data, ensuring the transfer of information between the data space and the parameter space. Experimental results on the CIFAR-10 and FashionMNIST datasets confirm that the resulting vector representations allow for a high degree of accuracy in distinguishing between models trained on different data subsets, without resorting to complex vector generation algorithms or using the original samples. These results suggest that the parameters of trained autoencoders can be viewed as sample representations.

自编码器参数表示数据指纹

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。