arXiv:2505.20333cs.CLcs.AI2025-05被引 1

提出多尺度流形对齐框架,揭示大模型内部信息层级结构。

Multi-Scale Manifold Alignment for Interpreting Large Language Models: A Unified Information-Geometric Framework

  • 将大模型表征分解为局部、中间、全局流形,学习跨尺度几何对齐
  • 在多个模型上实现显著的相对KL下降与互信息提升(统计显著)
  • 可实现分尺度干预,适用于模型可解释性与可控生成研究

我们提出多尺度流形对齐(MSMA)框架,将大语言模型(LLM)的表示分解为局部、中间和全局流形,并学习保持几何与信息一致性的跨尺度映射。在GPT-2、BERT、RoBERTa和T5等多个模型中,我们观察到一致的层次化模式,且MSMA在多种估计器下均显著提升对齐指标(如相对KL减少和互信息增益,种子间具有统计显著性)。在不同尺度上进行受控干预,会带来词汇多样性、句子结构和话语连贯性上不同的、与架构相关的效应。尽管理论分析基于理想化假设,但实证结果表明,多目标对齐提供了一个实用的视角,用于分析跨尺度信息流动并指导表示层面的控制。

原文摘要 · Abstract (English)

We present Multi-Scale Manifold Alignment(MSMA), an information-geometric framework that decomposes LLM representations into local, intermediate, and global manifolds and learns cross-scale mappings that preserve geometry and information. Across GPT-2, BERT, RoBERTa, and T5, we observe consistent hierarchical patterns and find that MSMA improves alignment metrics under multiple estimators (e.g., relative KL reduction and MI gains with statistical significance across seeds). Controlled interventions at different scales yield distinct and architecture-dependent effects on lexical diversity, sentence structure, and discourse coherence. While our theoretical analysis relies on idealized assumptions, the empirical results suggest that multi-objective alignment offers a practical lens for analyzing cross-scale information flow and guiding representation-level control.

大模型解释流形对齐信息几何可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。