用预训练大模型构建可分析的文本压缩表示,提升语义解耦性。
LangVAE and LangSpace: Building and Probing for Language Model VAEs
- 基于大语言模型构建可模块化组合的变分自编码器框架。
- 不同架构组合下展现显著的泛化与解耦性能差异。
- 配套探针工具支持向量插值、聚类可视化等分析方法。
我们提出 LangVAE,一种在预训练大语言模型基础上构建变分自编码器(VAEs)的新框架。该语言模型 VAE 可将预训练组件的知识压缩为更紧凑且语义解耦的表示。通过配套框架 LangSpace,可实现向量遍历、插值、解耦度量及聚类可视化等探针分析。LangVAE 与 LangSpace 提供灵活、高效、可扩展的文本表示构建与分析方式,支持 HuggingFace Hub 上现有模型的快速集成。我们还测试了多种编码器-解码器组合及标注输入,揭示了不同架构家族和规模间在泛化与解耦方面的广泛交互关系。结果表明,该框架有望系统化推进文本表示的实验与理解。
原文摘要 · Abstract (English)
We present LangVAE, a novel framework for modular construction of variational autoencoders (VAEs) on top of pre-trained large language models (LLMs). Such language model VAEs can encode the knowledge of their pre-trained components into more compact and semantically disentangled representations. The representations obtained in this way can be analysed with the LangVAE companion framework: LangSpace, which implements a collection of probing methods, such as vector traversal and interpolation, disentanglement measures, and cluster visualisations. LangVAE and LangSpace offer a flexible, efficient and scalable way of building and analysing textual representations, with simple integration for models available on the HuggingFace Hub. Additionally, we conducted a set of experiments with different encoder and decoder combinations, as well as annotated inputs, revealing a wide range of interactions across architectural families and sizes w.r.t. generalisation and disentanglement. Our findings demonstrate a promising framework for systematising the experimentation and understanding of textual representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。