arXiv:2601.13851cs.LGstat.ML2026-01

用距离反推输入,让SOM能生成清晰可控的图像过渡。

Inverting Self-Organizing Maps: A Unified Activation-Based Framework

  • 基于原型距离反解输入,利用欧氏几何唯一性原理。
  • 生成轨迹保持高分类置信度,比线性插值更清晰锐利。
  • 无需采样或训练,适合需要可解释生成的场景。

自组织映射(SOM)能保持高维数据拓扑结构,但其作为生成模型的应用仍不充分。我们发现,SOM的激活模式——即到原型的平方距离——可通过经典欧氏距离几何结论逆向恢复原始输入:一个D维空间中的点可由其到D+1个仿射独立参考点的距离唯一确定。我们推导出对应的线性系统,并刻画了逆问题适定的条件。在此基础上,提出曼达福德感知统一SOM反演与控制(MUSIC)更新规则,通过修改特定原型的平方距离而保持其余不变,生成符合SOM分段线性结构的、语义连贯的可控轨迹。引入Tikhonov正则化稳定更新过程,确保高维空间中平滑运动。与变分或扩散生成模型不同,MUSIC无需采样、隐式先验或学习解码器,完全基于原型几何操作。若无扰动,可精确重构输入;指定目标原型或簇时,能生成一致的语义过渡。我们在合成高斯混合、MNIST数字和LFW人脸数据集上验证该框架,结果表明,所有设置下,MUSIC轨迹均保持高分类置信度,中间图像显著比线性插值更清晰,并揭示了学习映射的可解释几何结构。

原文摘要 · Abstract (English)

Self-Organizing Maps (SOMs) provide topology-preserving projections of high-dimensional data, yet their use as generative models remains largely unexplored. We show that the activation pattern of a SOM -- the squared distances to its prototypes -- can be \emph{inverted} to recover the exact input, following from a classical result in Euclidean distance geometry: a point in $D$ dimensions is uniquely determined by its distances to $D{+}1$ affinely independent references. We derive the corresponding linear system and characterize the conditions under which inversion is well-posed. Building on this mechanism, we introduce the \emph{Manifold-Aware Unified SOM Inversion and Control} (MUSIC) update rule, which modifies squared distances to selected prototypes while preserving others, producing controlled, semantically meaningful trajectories aligned with the SOM's piecewise-linear structure. Tikhonov regularization stabilizes the update and ensures smooth motion in high dimensions. Unlike variational or diffusion-based generative models, MUSIC requires no sampling, latent priors, or learned decoders: it operates entirely on prototype geometry. If no perturbation is applied, inversion recovers the exact input; when a target prototype or cluster is specified, MUSIC produces coherent semantic transitions. We validate the framework on synthetic Gaussian mixtures, MNIST digits, and the Labeled Faces in the Wild dataset. Across all settings, MUSIC trajectories maintain high classifier confidence, produce significantly sharper intermediate images than linear interpolation, and reveal an interpretable geometric structure of the learned map.

生成模型SOM几何生成可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。