探究德语BERT如何理解复合词语义,发现其效果弱于英语模型。
Probing BERT for German Compound Semantics
- 通过变换词元、层和大小写,测试BERT对德语复合词的语义编码能力。
- 早期层中可恢复复合性信息,但整体准确率显著低于英语相关研究。
- 适合关注多语言NLP差异与德语语言特性的研究人员。
本文研究预训练德语BERT在多大程度上编码了名词复合词的语义知识。我们系统地变化目标词元、网络层数以及大小写敏感与否的模型配置,并通过预测868个标准复合词的组合性来评估性能。分析Transformer架构中的表示模式后,观察到的趋势与先前针对英语的研究类似:组合性信息在早期层中最易恢复。然而,最强结果仍明显落后于英语研究报道的水平,表明德语中的该任务更具挑战性。这可能源于德语中构词的更高产出性,以及由此带来的成分层面歧义增加,包括在本研究的目标复合词集中。
原文摘要 · Abstract (English)
This paper investigates the extent to which pretrained German BERT encodes knowledge of noun compound semantics. We comprehensively vary combinations of target tokens, layers, and cased vs. uncased models, and evaluate them by predicting the compositionality of 868 gold standard compounds. Looking at representational patterns within the transformer architecture, we observe trends comparable to equivalent prior work on English, with compositionality information most easily recoverable in the early layers. However, our strongest results clearly lag behind those reported for English, suggesting an inherently more difficult task in German. This may be due to the higher productivity of compounding in German than in English and the associated increase in constituent-level ambiguity, including in our target compound set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。