通过张量场收敛提升语言模型表示的一致性
Statistical Coherence Alignment for Large Language Model Representation Learning Through Tensor Field Convergence
- 利用张量场收敛约束词元表示,强化统计依赖关系
- 降低困惑度,提升分类准确率,改善罕见词嵌入
- 适合需要高上下文一致性的生成任务研究者
表示学习在构建内部嵌入以捕捉语言统计特性方面起核心作用,影响生成文本的连贯性和上下文一致性。本文提出统计连贯性对齐方法,通过张量场收敛机制引导嵌入反映语言数据中的固有统计依赖关系。建立了量化连贯性对齐的数学框架,包含一个优化训练迭代中表示一致性的损失函数。实证评估表明,引入连贯性约束可降低困惑度、提升分类准确率并改进稀有词嵌入,促进更稳定的表示空间。与基线模型的对比分析显示,该方法增强了内部结构可解释性,使嵌入在保留上下文依赖的同时缓解表示坍缩。对连贯性得分分布的影响表明,对齐机制强化了跨多样语言结构的语义完整性,实现更均衡的嵌入组织。计算评估指出,尽管方法带来额外内存和训练开销,但结构化优化过程在需要高上下文保真度的应用中值得权衡。实验结果验证了连贯性对齐在优化词元表示方面的有效性,为利用统计依赖改进语言模型训练提供了新思路。
原文摘要 · Abstract (English)
Representation learning plays a central role in structuring internal embeddings to capture the statistical properties of language, influencing the coherence and contextual consistency of generated text. Statistical Coherence Alignment is introduced as a method to enforce structured token representations through tensor field convergence, guiding embeddings to reflect statistical dependencies inherent in linguistic data. A mathematical framework is established to quantify coherence alignment, integrating a loss function that optimizes representational consistency across training iterations. Empirical evaluations demonstrate that applying coherence constraints improves perplexity, enhances classification accuracy, and refines rare word embeddings, contributing to a more stable representation space. Comparative analyses with baseline models reveal that the proposed method fosters a more interpretable internal structure, ensuring that embeddings retain contextual dependencies while mitigating representation collapse. The impact on coherence score distributions suggests that the alignment mechanism strengthens semantic integrity across diverse linguistic constructs, leading to a more balanced organization of learned embeddings. Computational assessments indicate that while the method introduces additional memory and training costs, the structured optimization process justifies the trade-offs in applications requiring heightened contextual fidelity. Experimental results validate the effectiveness of coherence alignment in optimizing token representations, providing insights into how statistical dependencies can be leveraged to improve language model training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。