用拓扑方法追踪大模型对齐过程中的内部表征变化。
Tracking Representation Dynamics in Large Language Models with Persistent Homology

- 用持久同调分析激活空间的拓扑结构演化。
- 早期训练阶段完成大部分拓扑重构,随后快速稳定。
- 不同对齐目标引发可区分的拓扑轨迹,适合研究模型内在机制。
大型语言模型通常通过监督微调进行对齐,但其内部表征在这一过程中如何演化仍不明确。本文利用持久同调技术,追踪微调期间激活空间的拓扑结构变化。在四个参数量从10亿到70亿的Transformer语言模型上,针对有益、无害及混合三种对齐目标进行分析,发现多数拓扑重组集中在训练初期。密集检查点分析显示,拓扑活动出现瞬时峰值后迅速趋于稳定。此外,不同对齐目标导致可区分的拓扑轨迹,而指令微调与预训练模型表现出本质不同的演化模式。结果表明,持久同调为理解对齐过程提供了补充视角,揭示了行为指标之外的表征级变化。
原文摘要 · Abstract (English)
Large language models are commonly aligned through supervised fine-tuning, yet little is known about how their internal representations evolve during this process. We study alignment dynamics using persistent homology by tracking the topology of activation spaces throughout fine-tuning. Across four transformer language models ranging from 1B to 7B parameters and three alignment objectives corresponding to helpful, harmless, and mixed training data, we find that the majority of topological reorganization occurs during the earliest stages of training. A dense checkpoint analysis reveals a transient peak in topological activity followed by rapid stabilization. We further show that different alignment objectives induce distinguishable topological trajectories, while instruction-tuned and pretrained models exhibit qualitatively different patterns of evolution. Our results suggest that persistent homology provides a complementary perspective on alignment, revealing representation-level changes that are not apparent from behavioral metrics alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。