用提示调优让大模型帮孟加拉语语法错误自动解释,效果接近人工。
Leveraging Prompt-Tuning for Bengali Grammatical Error Explanation Using Large Language Models
- 设计三步提示调优流程:识别错误、修正句子、生成解释。
- GPT-4在自动评估中F1提升5.26%,准确匹配率提高6.95%。
- 适合需要低资源语言语法纠错的教育或NLP研究者使用。
我们提出一种基于先进大语言模型(如GPT-4、GPT-3.5 Turbo、Llama-2-70b)的三步提示调优方法,用于孟加拉语语法错误解释(BGEE)。该方法包括识别并分类孟加拉语句子中的语法错误,生成修正后的句子,并为每类错误提供自然语言解释。通过自动化指标和母语专家的人工评估对系统性能进行测试。结果显示,表现最佳的GPT-4相较基线模型在自动评估中F1分数提升5.26%,精确匹配率提高6.95%;错误类型识别错误率下降25.51%,错误解释错误率下降26.27%。尽管如此,其表现仍不及人类基准。
原文摘要 · Abstract (English)
We propose a novel three-step prompt-tuning method for Bengali Grammatical Error Explanation (BGEE) using state-of-the-art large language models (LLMs) such as GPT-4, GPT-3.5 Turbo, and Llama-2-70b. Our approach involves identifying and categorizing grammatical errors in Bengali sentences, generating corrected versions of the sentences, and providing natural language explanations for each identified error. We evaluate the performance of our BGEE system using both automated evaluation metrics and human evaluation conducted by experienced Bengali language experts. Our proposed prompt-tuning approach shows that GPT-4, the best performing LLM, surpasses the baseline model in automated evaluation metrics, with a 5.26% improvement in F1 score and a 6.95% improvement in exact match. Furthermore, compared to the previous baseline, GPT-4 demonstrates a decrease of 25.51% in wrong error type and a decrease of 26.27% in wrong error explanation. However, the results still lag behind the human baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。