用大模型辅助粤语爱沙尼亚语语法工程,发现生成效果有限但有参考价值。
Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram Treebanks
- 在平行语法项目中构建粤语与爱沙尼亚语树库,保持抽象功能层面的平行性。
- 大模型翻译表现差,生成句法结构虽部分合理但缺乏跨语言抽象能力。
- 生成结果可作分析参考,适合需人工验证的多语言语法研究者使用。
语法工程需要语言形式化与计算实现的双重专长,尤其在平衡跨语言一致性与语言特异性方面。本文介绍了在平行语法(ParGram)项目下构建粤语与爱沙尼亚语树库的过程,确保抽象功能层面的平行性。我们进一步探讨了多语言大模型在语法工程中的方法论潜力与局限,聚焦于粤语-爱沙尼亚语翻译及使用OpenAI的gpt-oss-120b模型生成形式句法结构。结果显示,翻译性能普遍不佳,且不受提示语言影响;句法结构生成虽产生部分有意义输出,但在需要跨语言抽象的任务上表现较差。然而,大模型生成的内容仍可能提供参考价值,提示替代分析并(部分)捕捉谓词-论元关系。总体而言,研究揭示了大模型在协作语法工程中的潜力与局限,强调专家主导分析与验证的持续重要性。
原文摘要 · Abstract (English)
Grammar engineering requires expertise in linguistic formalism and computational implementation, especially in parallel grammar projects that balance cross-linguistic consistency with language-specific properties. This paper presents the development of Cantonese and Irish treebanks within the Parallel Grammar (ParGram) Project, where linguistic parallelism is maintained at an abstract functional level. We also investigate the methodological potential and limitations of using multilingual LLMs to support grammar engineering, focusing on Cantonese-Irish translation and the generation of formal syntactic structures using OpenAI's gpt-oss-120b model. The results show that translation performance was generally unsatisfactory and unaffected by prompt language. For syntactic structure generation, the model produced some structurally meaningful outputs, but performed poorly on tasks requiring cross-linguistic abstraction. Nonetheless, LLM-generated outputs may still offer some reference value by suggesting alternative analyses and (partially) capturing predicate-argument relations. Overall, our findings highlight both the potential and limitations of using LLMs in collaborative grammar engineering, while underscoring the continued importance of expert-driven analysis and verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。