arXiv:2506.20083cs.CL2025-06综述

用自编码器构建语义空间几何,连接符号与分布语义

Bridging Compositional and Distributional Semantics: A Survey on Latent Semantic Geometry via AutoEncoder

  • 通过自编码器框架分析语义空间的几何结构
  • 对比VAE、VQVAE、SAE三种模型的语义可解释性差异
  • 适合关注语言模型可解释性的研究者阅读

将组合性与符号属性融入当前的分布语义空间,可提升基于Transformer的自回归语言模型在可解释性、可控性、组合能力及泛化性能方面的表现。本文从组合语义视角出发,提出一种新的潜在空间几何理解路径,称为语义表征学习,旨在弥合符号语义与分布语义之间的鸿沟。我们综述并比较了三种主流自编码器架构:变分自编码器(VAE)、向量量化自编码器(VQVAE)和稀疏自编码器(SAE),分析它们在语义结构与可解释性方面所诱导的独特潜在几何特征。

原文摘要 · Abstract (English)

Integrating compositional and symbolic properties into current distributional semantic spaces can enhance the interpretability, controllability, compositionality, and generalisation capabilities of Transformer-based auto-regressive language models (LMs). In this survey, we offer a novel perspective on latent space geometry through the lens of compositional semantics, a direction we refer to as \textit{semantic representation learning}. This direction enables a bridge between symbolic and distributional semantics, helping to mitigate the gap between them. We review and compare three mainstream autoencoder architectures-Variational AutoEncoder (VAE), Vector Quantised VAE (VQVAE), and Sparse AutoEncoder (SAE)-and examine the distinctive latent geometries they induce in relation to semantic structure and interpretability.

语义表征自编码器可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。