多任务训练让模型世界表征收敛,但某些任务反而阻碍新实体融入。
Convergent World Representations and Divergent Tasks
- 通过5075个城市坐标与7种几何任务构建可控实验环境。
- 多任务训练使不同任务的表征几何趋于一致,验证了统一表征假设。
- 部分任务会破坏新城市信息的整合能力,影响模型泛化性能。
尽管神经表征在现代深度学习中至关重要,但其几何结构及对下游适应性的影响机制仍不清晰。本文构建了一个框架,明确分离世界本身、数据生成过程与模型表征,以在受控环境下研究该问题。使用5,075个城市坐标定义世界,并通过7种几何任务生成自回归训练数据。结果表明,不同任务导致质与量上均不同的世界表征几何。然而,多任务训练促使表征收敛:即使任务不重叠,模型仍发展出对齐的几何结构,为柏拉图式统一表征假说中的多任务扩展假说提供了控制性证据。为进一步研究适应性,我们先在所有任务上预训练模型,再测试通过微调将新城市(新实体)整合进表征空间的能力。令人惊讶的是,尽管经过多任务预训练,某些任务(称为发散任务)仍会显著损害新实体的表征整合,降低泛化性能。研究揭示:多关系任务训练能可靠产生收敛的世界表征,但隐藏的发散任务可能通过微调严重破坏新实体的集成能力。
原文摘要 · Abstract (English)
While neural representations are central to modern deep learning, the conditions governing their geometry and their roles in downstream adaptability remain poorly understood. We develop a framework clearly separating the underlying world, the data generation process and the resulting model representations to study these questions in a controlled setup. 5,075 city coordinates define the world and 7 geometric tasks generate the training data for autoregressive training. We find that different tasks give rise to qualitatively and quantitatively distinct world representation geometries. However, multi-task training drives convergence of world representations: models trained on non-overlapping tasks develop aligned geometric representations, providing controlled evidence for the Multitask Scaling Hypothesis of the Platonic Representation Hypothesis. To study adaptation, we pretrain models on all tasks, then test whether new entities (cities) can be consistently integrated into the representation space via fine-tuning. Surprisingly, we find that despite multi-task pretraining, some tasks, which we call divergent, actively harm the representational integration of new entities and harm generalization. Our results show that training on multiple relational tasks reliably produces convergent world representations, but lurking divergent tasks can catastrophically harm new entity integration via fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。