arXiv:2607.07047cs.CLcs.AI2026-07

用黎曼几何分析语言模型嵌入,提升分类效果并验证其可解释性。

Riemannian Geometry for Pre-trained Language Model Embeddings

论文配图:Riemannian Geometry for Pre-trained Language Model Embeddings
图 1 · 摘自论文原文
  • 通过提取编码器雅可比矩阵的逐标记拉回度量,结合对称正定流形上的弗雷歇均值聚合。
  • 在CoLA、CREAK、RTE上优于欧氏均值池化,但在无词法噪声的FEVER-Symmetric上表现随机。
  • 几何聚合本身即具优势,训练编码器仅在知识密集型数据集上带来额外增益。

理解预训练语言模型嵌入的几何结构对可解释性和安全性至关重要。本文探讨句级分类信号是否存在于上下文标记嵌入的黎曼几何中,通过从学习编码器的解析雅可比矩阵中提取逐标记拉回度量,并在对称正定(SPD)流形上使用弗雷歇均值聚合,提出黎曼均值池化(RMP)方法。在具有非平凡语言结构的三个数据集(CoLA、CREAK、RTE)上,RMP表现优于欧氏均值池化;而在去除标注驱动词法伪影的FEVER-Symmetric基准上,该方法正确维持在随机水平。消融实验表明,随机初始化编码器配合弗雷歇聚合已在两个信号数据集上超越欧氏池化,说明性能提升主要源于几何聚合机制而非学习到的流形结构;训练编码器仅在最依赖知识的CREAK数据集上贡献额外信号。

原文摘要 · Abstract (English)

Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety. We ask whether sentence-level classification signal lives in the Riemannian geometry of contextual token embeddings, and probe it by extracting per-token pullback metrics from a learned encoder's analytical Jacobian and aggregating them with the Fréchet mean on the symmetric positive definite (SPD) manifold; we call this procedure Riemannian Mean Pooling (RMP). Across three datasets with non-trivial linguistic structure (CoLA, CREAK, RTE), RMP outperforms Euclidean mean pooling, while on FEVER-Symmetric, a benchmark constructed to remove annotation-driven lexical artifacts, the method correctly stays at chance. Ablations show that a randomly initialised encoder combined with Fréchet aggregation already beats Euclidean pooling on two of the three signal-bearing datasets, localising the source of the gain to the geometric aggregation rather than to learned manifold structure; the trained encoder contributes additional signal specifically on CREAK, the most knowledge-heavy of the three signal-bearing datasets.

语言模型黎曼几何嵌入分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。