arXiv:2511.20120cs.CLcs.AI2025-11中稿 · -demonstration at …

用提示工程让大模型在数据少时也能高效纠正印地语等语言的语法错误。

"When Data is Scarce, Prompt Smarter"... Approaches to Grammatical Error Correction in Low-Resource Settings

  • 用少量示例和精心设计的提示,让大模型直接适应低资源语言
  • 在泰米尔、印地语等五种印地语系语言上取得领先效果
  • 适合资源匮乏语言的语法纠错研究者参考

语法错误纠正(GEC)是自然语言处理中的重要任务,旨在自动检测并修正文本中的语法错误。尽管基于Transformer的模型和大规模标注数据显著提升了英语等高资源语言的性能,但大多数印地语系语言仍面临挑战,受限于资源稀缺、语言多样性及复杂形态学。本文探索使用前沿大语言模型(如GPT-4.1、Gemini-2.5、LLaMA-4)结合少量示例策略的提示方法,以适应低资源环境。实验表明,即使基础提示策略(零样本、少样本)也使这些大模型显著优于微调后的印地语模型(如Sarvam-22B),凸显当代大模型在多语言GEC中的卓越泛化能力。通过精心设计的提示与轻量级适配,多种印地语系语言的纠错质量显著提升。在共享任务中,我们取得了领先成绩:泰米尔语(GLEU: 91.57)第1名,印地语(GLEU: 85.69)第1名,泰卢固语(GLEU: 85.22)第2名,孟加拉语(GLEU: 92.86)第4名,马拉雅拉姆语(GLEU: 92.97)第5名。结果表明提示驱动技术有效,大模型具备弥合多语言资源差距的潜力。

原文摘要 · Abstract (English)

Grammatical error correction (GEC) is an important task in Natural Language Processing that aims to automatically detect and correct grammatical mistakes in text. While recent advances in transformer-based models and large annotated datasets have greatly improved GEC performance for high-resource languages such as English, the progress has not extended equally. For most Indic languages, GEC remains a challenging task due to limited resources, linguistic diversity and complex morphology. In this work, we explore prompting-based approaches using state-of-the-art large language models (LLMs), such as GPT-4.1, Gemini-2.5 and LLaMA-4, combined with few-shot strategy to adapt them to low-resource settings. We observe that even basic prompting strategies, such as zero-shot and few-shot approaches, enable these LLMs to substantially outperform fine-tuned Indic-language models like Sarvam-22B, thereby illustrating the exceptional multilingual generalization capabilities of contemporary LLMs for GEC. Our experiments show that carefully designed prompts and lightweight adaptation significantly enhance correction quality across multiple Indic languages. We achieved leading results in the shared task--ranking 1st in Tamil (GLEU: 91.57) and Hindi (GLEU: 85.69), 2nd in Telugu (GLEU: 85.22), 4th in Bangla (GLEU: 92.86), and 5th in Malayalam (GLEU: 92.97). These findings highlight the effectiveness of prompt-driven NLP techniques and underscore the potential of large-scale LLMs to bridge resource gaps in multilingual GEC.

语法纠错大模型低资源提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。