用语义单元纠缠编码压缩大模型,兼顾性能与效率
SEE: Sememe Entanglement Encoding for Transformer-bases Models Compression
- 基于语义单元的低秩近似与量子纠缠重构
- 模型参数和计算量显著降低,性能稳定
- 适合资源受限场景下的大模型部署
基于Transformer的大语言模型虽具突破性能力,但存储与计算成本过高,限制了其在资源受限场景的应用。为此,我们提出语义纠缠编码(Sememe Entanglement Encoding, SEE)算法,借助专家先验知识,通过低秩近似实现模型压缩。在纠缠嵌入中,基本语义单元(如语义素)被表示为低维向量,并通过广义量子纠缠组合重构为高维词嵌入。该方法适用于不同规模的Transformer模型。实验表明,该方法在显著压缩模型参数与计算开销的同时保持性能稳定。
原文摘要 · Abstract (English)
Transformer-based large language models exhibit groundbreaking capabilities, but their storage and computational costs are prohibitively high, limiting their application in resource-constrained scenarios. An effective approach is to eliminate redundant model parameters and computational costs while incorporating efficient expert-derived knowledge structures to achieve a balance between compression and performance. Therefore, we propose the \textit{Sememe Entanglement Encoding (SEE)} algorithm. Guided by expert prior knowledge, the model is compressed through the low-rank approximation idea. In Entanglement Embedding, basic semantic units such as sememes are represented as low-dimensional vectors, and then reconstructed into high-dimensional word embeddings through the combination of generalized quantum entanglement. We adapt the Sememe Entanglement Encoding algorithm to transformer-based models of different magnitudes. Experimental results indicate that our approach achieves stable performance while compressing model parameters and computational costs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。