发现语法归纳模型的分布坍缩问题,提出新方法让模型更简洁高效
Probability Distribution Collapse: A Critical Bottleneck to Compact Unsupervised Neural Grammar Induction
- 通过分析神经参数化中的分布坍缩机制,定位瓶颈根源
- 在多语言数据上实现更小模型、更高解析性能
- 适合研究语法学习与模型压缩的学者参考
无监督神经语法归纳旨在从语言数据中学习可解释的层次结构。然而,现有模型存在表达能力瓶颈,常导致过于庞大却表现不佳的语法结构。本文识别出核心问题——概率分布坍缩,是这一局限的根本原因。我们分析了该现象在神经参数化关键组件中出现的时机与方式,并提出针对性解决方案:坍缩缓解型神经参数化。实验表明,该方法显著提升解析性能,同时支持在多种语言下使用更紧凑的语法结构。
原文摘要 · Abstract (English)
Unsupervised neural grammar induction aims to learn interpretable hierarchical structures from language data. However, existing models face an expressiveness bottleneck, often resulting in unnecessarily large yet underperforming grammars. We identify a core issue, $\textit{probability distribution collapse}$, as the underlying cause of this limitation. We analyze when and how the collapse emerges across key components of neural parameterization and introduce a targeted solution, $\textit{collapse-relaxing neural parameterization}$, to mitigate it. Our approach substantially improves parsing performance while enabling the use of significantly more compact grammars across a wide range of languages, as demonstrated through extensive empirical analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。