arXiv:2502.00246cs.CL2025-02

通过动态重排权重张量,提升大模型长程依赖建模能力。

Context-Preserving Tensorial Reconfiguration in Large Language Model Training

  • 采用结构化分解与自适应压缩,动态重构模型权重张量。
  • 在长序列任务中降低困惑度,提升召回准确率,内存消耗减少。
  • 适合需要高效长上下文理解的模型优化场景。

神经网络处理长距离依赖始终受限于计算瓶颈和低效的上下文保留机制。张量操作为重构模型表示提供了基础,但传统架构难以在不引入过多复杂性的情况下应用此类技术。本文提出一种新方法——上下文保持张量重排(CPTR),通过结构化分解与自适应收缩实现权重张量的动态重组,显著增强上下文整合能力,且计算开销可控。实证表明,CPTR增强了长序列中的连贯性保留,使困惑度下降,长上下文任务的召回准确率提升。相比基线模型,其在保持语言生成流畅性与准确性的同时,具备更高的计算效率和更低的内存占用。梯度稳定性指标显示权重更新方差更小,训练更稳定。对比研究证实,张量重排有助于构建更稳定、高效的语言建模架构。结果表明,CPTR在需要长程上下文理解与高效内存利用的任务中具有显著潜力。

原文摘要 · Abstract (English)

Handling long-range dependencies in neural architectures has remained a persistent challenge due to computational limitations and inefficient contextual retention mechanisms. Tensorial operations have provided a foundation for restructuring model representations, yet conventional architectures have struggled to incorporate such techniques without introducing excessive complexity. A novel approach, Context-Preserving Tensorial Reconfiguration (CPTR), enables dynamic reorganization of weight tensors through structured factorization and adaptive contraction, allowing for enhanced contextual integration without substantial computational overhead. Empirical evaluations demonstrate that CPTR improves coherence retention across extended sequences, leading to measurable reductions in perplexity and improved recall accuracy for long-context tasks. Performance comparisons reveal that CPTR-enhanced models exhibit greater computational efficiency and reduced memory consumption while maintaining competitive language generation fluency and accuracy. Gradient stability metrics further validate the improved training efficiency, revealing more controlled variance in weight updates. Comparative studies across baseline and CPTR-enhanced models confirm that tensorial reconfiguration contributes to more stable and computationally efficient language modeling. The findings support the potential of CPTR in refining contemporary neural architectures for tasks requiring long-range contextual understanding and efficient memory utilization.

张量重排长上下文模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。