arXiv:2511.15260cs.CL2025-11被引 1

小模型在印地语语法纠错中表现超群,揭示数据与评测标准的关键问题。

IndicGEC: Powerful Models, or a Measurement Mirage?

  • 用小模型零样本提示实现高效纠错,无需大量训练数据。
  • 印地语和泰卢固语得分分别达84.31和83.78,排名第二和第四。
  • 揭示印地语语法纠错任务中数据质量与评估指标的潜在缺陷。

本文报告了团队NRC在BHASHA-Task 1语法错误纠正共享任务中的结果,涵盖5种印度语言。通过针对不同规模的语言模型(40亿到大型专有模型)进行零/少样本提示,我们在泰卢固语中取得第4名,印地语第2名,GLEU得分分别为83.78和84.31。本研究进一步扩展至另外三种语言——泰米尔语、马拉雅拉姆语和孟加拉语,并深入分析了数据质量和评估指标。结果表明,小型语言模型具有显著潜力,同时强调构建高质量数据集及适用于印度文字脚本的合理评估指标的重要性。

原文摘要 · Abstract (English)

In this paper, we report the results of the TeamNRC's participation in the BHASHA-Task 1 Grammatical Error Correction shared task https://github.com/BHASHA-Workshop/IndicGEC2025/ for 5 Indian languages. Our approach, focusing on zero/few-shot prompting of language models of varying sizes (4B to large proprietary models) achieved a Rank 4 in Telugu and Rank 2 in Hindi with GLEU scores of 83.78 and 84.31 respectively. In this paper, we extend the experiments to the other three languages of the shared task - Tamil, Malayalam and Bangla, and take a closer look at the data quality and evaluation metric used. Our results primarily highlight the potential of small language models, and summarize the concerns related to creating good quality datasets and appropriate metrics for this task that are suitable for Indian language scripts.

语法纠错小模型数据质量多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。