用单纯形几何提升模型隐空间各向同性,无需重训练
Shrink the longest: improving latent space isotropy with symplicial geometry
- 基于单纯形几何的正则化方法,利用流形拓扑结构优化隐空间分布
- 显著降低微调过程中的各向异性,下游任务性能提升明显
- 不增加推理开销,不需额外数据,适合实际部署场景
尽管基于Transformer的模型在深度学习领域占据主导地位,但多项研究发现其嵌入空间存在“表示退化问题”:嵌入值趋向于聚集在狭窄锥体内,导致隐空间高度各向异性。提高各向同性已被证实能提升静态与上下文语言模型的下游性能。然而,现有方法要么增加推理开销,要么需要大量数据进行模型重参数化。本文提出一种基于单纯形几何的新正则化技术,通过最大化使用Vietoris-Rips滤波从上下文嵌入中获得的条形码的持久熵,来改善隐表示的各向同性。实验表明,该方法在微调过程中显著降低各向异性,同时提升下游任务表现,且仅利用现有几何结构,无需重新参数化。
原文摘要 · Abstract (English)
Although transformer-based models have been dominating the field of deep learning, various studies of their embedding space have shown that they suffer from "representation degeneration problem": embeddings tend to be distributed in a narrow cone, making the latent space highly anisotropic. Increasing the isotropy has shown to improve performance in downstream tasks both in static and contextual language models. However, most of approaches either add inference overhead or require substantial amount of data for model reparametrization. We propose a novel regularization technique based on simplicial geometry to improve the isotropy of latent representations. The core idea of our method is based on maximizing the persistent entropy of barcodes obtained using Vietoris-Rips filtration from contextual embeddings in the underlying latent space. We demonstrate that the method leads to an increase in downstream performance while significantly lowering the anisotropy during fine-tuning by exploiting existing geometric structures instead of reparametrization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。