arXiv:2506.18120cs.CL2025-06中稿 · and published at L…被引 5

构建首个公开可用的英语句法可接受性数据集,助力语言学与机器学习研究。

The Syntactic Acceptability Dataset (Preview): A Resource for Machine Learning and Linguistic Analysis of English

  • 整合1000个英语句子,分自教材与《语言探究》期刊,覆盖当代语言学讨论。
  • 83%的句子在语法正确性与语感接受度上一致,'中间状态'频繁出现。
  • 机器学习模型预测语感优于预测语法,为自然语言处理提供新视角。

我们发布语法可接受性数据集的预览版本,该资源旨在支持句法与计算语言学研究。当前版本包含1000个英语语句,一半来自教材,一半来自期刊《语言探究》,以确保反映当代语言学话语。每条数据均标注其语法状态(依据句法形式系统)和可接受性状态(由母语者通过高标准众包评估获得)。尽管仍处于初步阶段,该数据集已是目前最大且公开可用的同类资源。初步分析揭示:语法正确性与语感接受度在约83%的情况下一致,'中间状态'普遍存在,符合既有研究;同时发现机器学习模型在预测可接受性方面表现显著优于预测语法正确性,此为新发现。未来工作将聚焦于扩展数据集规模。

原文摘要 · Abstract (English)

We present a preview of the Syntactic Acceptability Dataset, a resource being designed for both syntax and computational linguistics research. In its current form, the dataset comprises 1,000 English sequences from the syntactic discourse: Half from textbooks and half from the journal Linguistic Inquiry, the latter to ensure a representation of the contemporary discourse. Each entry is labeled with its grammatical status ("well-formedness" according to syntactic formalisms) extracted from the literature, as well as its acceptability status ("intuitive goodness" as determined by native speakers) obtained through crowdsourcing, with highest experimental standards. Even in its preliminary form, this dataset stands as the largest of its kind that is publicly accessible. We also offer preliminary analyses addressing three debates in linguistics and computational linguistics: We observe that grammaticality and acceptability judgments converge in about 83% of the cases and that "in-betweenness" occurs frequently. This corroborates existing research. We also find that while machine learning models struggle with predicting grammaticality, they perform considerably better in predicting acceptability. This is a novel finding. Future work will focus on expanding the dataset.

句法分析数据集机器学习语言学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。