用大模型自动生成伪标签,零样本实现语法能力精准评估
Zero-Shot Grammar Competency Estimation Using Large Language Model Generated Pseudo Labels
- 用大模型根据评分标准生成无标注数据的伪标签
- 在噪声标签下训练出高精度语法评分模型,准确率显著提升
- 适合资源有限的语法评估场景,无需人工标注
语法能力评估对书面与口语语言水平判断至关重要;但口语因自发性、无结构和不流畅性带来额外挑战。传统模型需大量专家标注,难以规模化。本文提出一种零样本语法能力评估框架,利用未标注数据与大语言模型(LLM)生成伪标签,无需人工标注。训练时通过基于评分标准的提示词,由LLM对未标注数据生成预测,作为伪标签,再通过新设计的抗噪声训练框架训练一个Transformer模型。实验表明,生成伪标签的LLM选择及训练中干净样本与噪声样本的比例,显著影响模型稳定性和准确性。定性分析验证了方法的鲁棒性与可解释性。结果证明该方法能高精度估计语法得分,为可扩展、低资源的语法评估系统奠定基础。
原文摘要 · Abstract (English)
Grammar competency estimation is essential for assessing linguistic proficiency in both written and spoken language; however, the spoken modality presents additional challenges due to its spontaneous, unstructured, and disfluent nature. Developing accurate grammar scoring models further requires extensive expert annotation, making large-scale data creation impractical. To address these limitations, we propose a zero-shot grammar competency estimation framework that leverages unlabeled data and Large Language Models (LLMs) without relying on manual labels. During training, we employ LLM-generated predictions on unlabeled data by using grammar competency rubric-based prompts. These predictions, treated as pseudo labels, are utilized to train a transformer-based model through a novel training framework designed to handle label noise effectively. We show that the choice of LLM for pseudo-label generation critically affects model performance and that the ratio of clean-to-noisy samples during training strongly influences stability and accuracy. Finally, a qualitative analysis of error intensity and score prediction confirms the robustness and interpretability of our approach. Experimental results demonstrate the efficacy of our approach in estimating grammar competency scores with high accuracy, paving the way for scalable, low-resource grammar assessment systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。