从语义标注语料中学习大规模可解释的构式语法
A Method for Learning Large-Scale Computational Construction Grammars from Semantically Annotated Corpora
- 基于句法结构和语义框架标注数据,自动构建构式语法网络
- 生成包含数万条构式的语法系统,支持开放域文本语义分析
- 适合研究英语论元结构的学者,助力构式语法规模化应用
我们提出一种从语言使用语料中学习大规模、广覆盖构式语法的方法。该方法以带有句法结构和语义框架标注的语句为输入,生成可解释的计算型构式语法,揭示句法结构与语义关系之间的复杂关联。所得语法由数十万条构式组成,形式化于流体构式语法(Fluid Construction Grammar)框架内。这些语法不仅能支持开放域文本的语义框架分析,还蕴含了训练数据中丰富的句法-语义使用模式信息。该方法及其生成的语法推动了基于使用的构式语言学方法的扩展,验证了若干核心构式语法假设的可扩展性,同时为在广覆盖语料中开展英语论元结构的构式主义研究提供了实用工具。
原文摘要 · Abstract (English)
We present a method for learning large-scale, broad-coverage construction grammars from corpora of language use. Starting from utterances annotated with constituency structure and semantic frames, the method facilitates the learning of human-interpretable computational construction grammars that capture the intricate relationship between syntactic structures and the semantic relations they express. The resulting grammars consist of networks of tens of thousands of constructions formalised within the Fluid Construction Grammar framework. Not only do these grammars support the frame-semantic analysis of open-domain text, they also house a trove of information about the syntactico-semantic usage patterns present in the data they were learnt from. The method and learnt grammars contribute to the scaling of usage-based, constructionist approaches to language, as they corroborate the scalability of a number of fundamental construction grammar conjectures while also providing a practical instrument for the constructionist study of English argument structure in broad-coverage corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。