arXiv:2603.23529cs.CLcs.AI2026-03

针对低资源语言康卡尼语,构建多文字体系指令微调模型并验证其效果。

Konkani LLM: Multi-Script Instruction Tuning and Evaluation for a Low-Resource Indian Language

  • 用Gemini 3生成10万条合成指令数据,覆盖德纳加里、罗马和卡纳达三种文字
  • 在机器翻译中,微调后的康卡尼LLM性能超越基础模型,并媲美甚至超过闭源模型
  • 发布多文字康卡尼基准测试集,支持跨文字评估,推动本地化研究

大型语言模型在低资源语言如康卡尼语上表现不佳,主要因训练数据极度稀缺且存在德纳加里、罗马和卡纳达三种书写系统。为此,我们提出基于Gemini 3生成的10万条合成指令微调数据集Konkani-Instruct-100k。通过评估Llama 3.1、Qwen2.5、Gemma 3等开源模型及闭源模型,建立严格基线。核心贡献包括开发一系列针对地区语言特征优化的康卡尼LLM微调模型,并正在构建多文字康卡尼基准测试集以支持跨文字语言评估。在机器翻译任务中,康卡尼LLM持续优于对应基础模型,在多个场景下达到甚至超越闭源模型表现。

原文摘要 · Abstract (English)

Large Language Models (LLMs) consistently under perform in low-resource linguistic contexts such as Konkani. This performance deficit stems from acute training data scarcity compounded by high script diversity across Devanagari, Romi and Kannada orthographies. To address this gap, we introduce Konkani-Instruct-100k, a comprehensive synthetic instruction-tuning dataset generated through Gemini 3. We establish rigorous baseline benchmarks by evaluating leading open-weights architectures including Llama 3.1, Qwen2.5 and Gemma 3 alongside proprietary closed-source models. Our primary contribution involves the development of Konkani LLM, a series of fine-tuned models optimized for regional nuances. Furthermore, we are developing the Multi-Script Konkani Benchmark to facilitate cross-script linguistic evaluation. In machine translation, Konkani LLM delivers consistent gains over the corresponding base models and is competitive with and in several settings surpasses proprietary baselines

低资源语言多文字指令微调康卡尼语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。