探究语言模型是否靠规则还是记忆掌握德语冠词
Understanding or Memorizing? A Case Study of German Definite Articles in Language Models
- 用梯度方法分析冠词变化的参数更新路径
- 发现不同性别-格组合间神经元影响高度重叠
- 结果支持记忆而非抽象规则的语法表征
语言模型在语法一致任务上表现良好,但其能力是基于规则泛化还是记忆存储仍不明确。本文针对德语单数定冠词(其形式取决于性别和格)进行研究,使用GRADIEND这一基于梯度的可解释性方法,学习特定性别-格组合下冠词转换对应的参数更新方向。结果显示,某一性别-格组合的更新常显著影响其他无关的性别-格设置,且受影响最显著的神经元在不同组合间存在大量重叠。该现象表明模型对德语定冠词的编码并非严格遵循规则,至少部分依赖于对具体搭配的记忆而非抽象语法规则。
原文摘要 · Abstract (English)
Language models perform well on grammatical agreement, but it is unclear whether this reflects rule-based generalization or memorization. We study this question for German definite singular articles, whose forms depend on gender and case. Using GRADIEND, a gradient-based interpretability method, we learn parameter update directions for gender-case specific article transitions. We find that updates learned for a specific gender-case article transition frequently affect unrelated gender-case settings, with substantial overlap among the most affected neurons across settings. These results argue against a strictly rule-based encoding of German definite articles, indicating that models at least partly rely on memorized associations rather than abstract grammatical rules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。