arXiv:2604.23985cs.AIcs.CL2026-04

发现语言模型的表示弯曲度影响预测不确定性,可调控生成行为。

Representational Curvature Modulates Behavioral Uncertainty in Large Language Models

论文配图:Representational Curvature Modulates Behavioral Uncertainty in Large Language Models
图 1 · 摘自论文原文
  • 用上下文曲率衡量表示轨迹弯曲程度,关联预测不确定性。
  • 曲率与下一词熵强相关,且随训练过程逐渐显现。
  • 调控曲率可改变生成不确定性,适合研究模型行为机制者。

在自回归大语言模型中,时间直线化解释了下一个词预测目标如何塑造表示。模型学习在各层间逐步拉直输入序列的表示轨迹,可能通过线性外推促进下一个词预测。然而,表示轨迹与词级行为之间的直接联系尚不明确。本文通过将上下文曲率——一个衡量近期上下文表示轨迹弯曲程度的几何量——与下一个词熵相关联,建立了这一联系。在 GPT-2 XL 和 Pythia-2.8B 两个模型中,上下文曲率与熵呈显著相关,且该关系在训练过程中逐步形成。扰动实验显示选择性依赖:通过沿轨迹方向干预可可靠调节熵,而几何错位的扰动则无效。最后,在训练中对表示进行更直线化的正则化,能适度降低词级熵,且不损害验证损失。结果表明,轨迹曲率是影响大语言模型行为不确定性的任务对齐表示特征。

原文摘要 · Abstract (English)

In autoregressive large language models (LLMs), temporal straightening offers an account of how the next-token prediction objective shapes representations. Models learn to progressively straighten the representational trajectory of input sequences across layers, potentially facilitating next-token prediction via linear extrapolation. However, a direct link between this trajectory and token-level behavior has been missing. We provide such a link by relating contextual curvature-a geometric measure of how sharply the representational trajectory bends over recent context-to next-token entropy. Across two models (GPT-2 XL and Pythia-2.8B), contextual curvature is correlated with entropy, and this relationship emerges during training. Perturbation experiments reveal selective dependence: manipulating curvature through trajectory-aligned interventions reliably modulates entropy, while geometrically misaligned perturbations have no effect. Finally, regularizing representations to be straighter during training modestly reduces token-level entropy without degrading validation loss. These results identify trajectory curvature as a task-aligned representational feature that influences behavioral uncertainty in LLMs.

语言模型表示曲率不确定性训练机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。