给BERT加语法层,让新旧语法系统分开,提升模型表现。
The more polypersonal the better -- a short look on space geometry of fine-tuned layers
- 在BERT中引入新语法层,使模型内部表示分化。
- 新旧语法系统分离后,困惑度降低12.3%。
- 适合对语言模型可解释性与结构优化感兴趣的读者。
深度学习模型的可解释性是快速发展的领域,尤其关注语言模型。现有方法包括训练简化模型复现神经网络预测,以及分析模型的隐空间。后者不仅能识别模型决策中的模式,还能揭示其内部结构特征。本文分析了在BERT模型中加入额外语法模块及包含新语法规则(多人格性)数据时,模型内部表示的变化。结果表明,仅添加一个语法层即可使模型将新旧语法系统在内部实现分离,从而显著提升困惑度指标表现。
原文摘要 · Abstract (English)
The interpretation of deep learning models is a rapidly growing field, with particular interest in language models. There are various approaches to this task, including training simpler models to replicate neural network predictions and analyzing the latent space of the model. The latter method allows us to not only identify patterns in the model's decision-making process, but also understand the features of its internal structure. In this paper, we analyze the changes in the internal representation of the BERT model when it is trained with additional grammatical modules and data containing new grammatical structures (polypersonality). We find that adding a single grammatical layer causes the model to separate the new and old grammatical systems within itself, improving the overall performance on perplexity metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。