改进零样本模型拼接的相对表示,提升稳定性和分类性能
Relative Representations: Topological and Geometric Perspectives
- 引入归一化使表示对非各向同性缩放和排列不变
- 用拓扑密集化损失增强类内聚类,提升零样本拼接效果
- 适合研究模型融合与零样本迁移的学者参考
相对表示是一种成熟的零样本模型拼接方法,通过深度神经网络隐空间的非训练变换实现。基于拓扑与几何洞察,本文提出两项改进:首先,在相对变换中引入归一化,使表示对非各向同性缩放和排列保持不变,后者与常见激活函数引起的参数空间对称性一致;其次,提出在微调相对表示时采用拓扑密集化,即一种鼓励类内聚类的拓扑正则化损失。我们在自然语言任务上进行了实证研究,结果表明两种改进均提升了零样本模型拼接的性能。
原文摘要 · Abstract (English)
Relative representations are an established approach to zero-shot model stitching, consisting of a non-trainable transformation of the latent space of a deep neural network. Based on insights of topological and geometric nature, we propose two improvements to relative representations. First, we introduce a normalization procedure in the relative transformation, resulting in invariance to non-isotropic rescalings and permutations. The latter coincides with the symmetries in parameter space induced by common activation functions. Second, we propose to deploy topological densification when fine-tuning relative representations, a topological regularization loss encouraging clustering within classes. We provide an empirical investigation on a natural language task, where both the proposed variations yield improved performance on zero-shot model stitching.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。