arXiv:2509.22598cs.CLcs.FL2025-09

证明了各类子正则语言可被线性模型完美区分,为自然语言建模提供可解释基础。

From Formal Language Theory to Statistical Learning: Finite Observability of Subregular Languages

  • 用决策谓词表示语言,发现所有子正则类均线性可分
  • 无噪声下合成实验实现完全可分,真实数据中特征契合语言约束
  • 适合关注可解释性与形式语言在语言建模中应用的研究者

我们证明,所有标准的子正则语言类在其决策谓词表示下均可线性分离。这确立了有限可观测性,并保证了使用简单线性模型的可学习性。合成实验在无噪声条件下验证了完全可分性;真实数据实验基于英语形态学表明,学习到的特征与已知的语言约束高度一致。结果表明,子正则层次结构为建模自然语言结构提供了严格且可解释的基础。真实数据实验所用代码已公开于 https://github.com/UTokyo-HayashiLab/subregular。

原文摘要 · Abstract (English)

We prove that all standard subregular language classes are linearly separable when represented by their deciding predicates. This establishes finite observability and guarantees learnability with simple linear models. Synthetic experiments confirm perfect separability under noise-free conditions, while real-data experiments on English morphology show that learned features align with well-known linguistic constraints. These results demonstrate that the subregular hierarchy provides a rigorous and interpretable foundation for modeling natural language structure. Our code used in real-data experiments is available at https://github.com/UTokyo-HayashiLab/subregular.

形式语言可学习性可解释性语言建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。