用粤语语法资源测试大模型生成语法能力,发现能辅助但难处理复杂约束。
How Useful are LLMs for Grammar Engineering? Cantonese ParGram Resources and Controlled Experimental Evaluation with English Baselines
- 以粤语和英语对照数据为标准,测试大模型从句子或目标结构生成语法
- GPT-5.4表现优于gpt-oss-120b,目标结构生成的语法更优
- 适合语言学专家做语法初稿辅助,需人工校验与修正
本文构建了新的粤语语法资源(Cantonese ParGram),并在受控实验中评估大模型在知识驱动型语法工程中的表现。以粤语资源为黄金标准,对比英文基线,研究OpenAI的gpt-oss-120b与GPT-5.4能否在不同提示条件下,从句子或目标形式结构中生成可机器处理的语法。结果表明,GPT-5.4优于gpt-oss-120b,且从目标形式结构生成的语法优于从句子生成。尽管模型能生成局部合理的短语结构规则、词项和模板,但在多构式交互约束下常出现协调失败。研究揭示当前大模型在集成进智能语法工作流时的能力与局限:可支持语法开发中间阶段,但语言分析、验证与精修仍需人类专家主导。同时,本研究贡献了新的粤语符号化语法资源。
原文摘要 · Abstract (English)
This paper presents new Cantonese ParGram resources and evaluates LLMs for knowledge-driven grammar engineering within a controlled experimental paradigm. Using Cantonese ParGram resources as gold standards, with corresponding English baselines, we investigate whether OpenAI's gpt-oss-120b and GPT-5.4 can generate machine-processable grammars from sentences and target formal structures under systematically varied prompting conditions. GPT-5.4 outperformed gpt-oss-120b, while grammars generated from target formal structures generally outperformed those generated from sentences. Although both models could generate locally plausible phrase-structure rules, lexical entries, and templates, they often struggled to coordinate interacting formal constraints, especially in multi-construction settings. The results characterize both the capabilities and limitations of current LLMs for potential integration into AI-assisted expert workflows: LLMs may support intermediate stages of grammar development, but human linguistic expertise remains central to analysis, validation, and refinement. The study also contributes new Cantonese symbolic grammatical resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。