arXiv:2510.21844cs.LGquant-ph2025-10

用量子张量网络压缩大模型,大幅降耗且保持精度。

KARIPAP: Quantum-Inspired Tensor Network Compression of Large Language Models Using Infinite Projected Entangled Pair States and Tensor Renormalization Group

  • 借量子物理思想,用iPEPS和TRG捕捉层间复杂关联。
  • 在LLaMA-2 7B上实现93%内存减少、70%参数压缩。
  • 适合追求高效部署与绿色计算的研究者。

大型语言模型(如ChatGPT和LLaMA)推动生成式AI快速发展,但其庞大的参数规模带来严重的计算与环境负担。高昂的训练成本、能耗及设备部署限制影响可及性。现有压缩方法(剪枝、蒸馏、低秩分解、量化)虽能减小规模,却忽视了层间复杂相关性。本文提出KARIPAP,一种基于量子张量网络的压缩框架,采用无限投影纠缠对态(iPEPS)与张量重整化群(TRG)收缩机制。相比一维矩阵乘积态,iPEPS可捕获注意力与深层Transformer中的多方向纠缠特性;TRG确保张量收缩在多项式时间内完成,使张量化成为可能,并保留关键相关性几何结构。在LLaMA-2 7B上的实验表明,该方法最多可减少93%内存占用、70%参数量,训练速度提升50%,推理速度加快25%,仅损失2-3%准确率。逐层纠缠分析揭示深层存在冗余,证实其适合张量分解。结果表明,现代大模型处于低维纠缠流形中,支持可扩展、节能且具备量子感知特性的智能架构设计。

原文摘要 · Abstract (English)

Large Language Models (LLMs) like ChatGPT and LLaMA drive rapid progress in generative AI, yet their huge parameter scales create severe computational and environmental burdens. High training costs, energy use, and limited device deployment hinder accessibility. Existing compression - pruning, distillation, low-rank, and quantization - reduces size but ignores complex inter-layer correlations. We propose KARIPAP, a quantum-inspired tensor network compression using Infinite Projected Entangled Pair States (iPEPS) and Tensor Renormalization Group (TRG) contraction. Unlike 1D Matrix Product States, iPEPS captures multi-directional entanglement in attention and deep transformer layers. TRG ensures polynomial-time contraction, making tensorization feasible while preserving key correlation geometry. Experiments on LLaMA-2 7B show up to 93% memory and 70% parameter reduction, with 50% faster training, 25% faster inference, and only 2-3% accuracy loss. Layer-wise entanglement profiling reveals redundancy in deeper layers, confirming their suitability for tensor factorization. KARIPAP demonstrates that modern LLMs occupy low-dimensional entanglement manifolds, enabling scalable, energy-efficient, and quantum-aware AI architectures.

大模型压缩张量网络量子启发高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。