arXiv:2510.04139cs.CLcs.LG2025-10

为小语种定制微调方法,让AI更好理解本土文化

Fine Tuning Methods for Low-resource Languages

  • 构建文化相关数据集,适配非英语语境
  • 在小语种上提升Gemma 2模型性能表现
  • 提供可复用的本地化AI开发路径

大型语言模型的发展并未惠及所有文化。这些模型主要基于英文文本和文化训练,导致在其他语言和文化背景下的表现不佳。本项目通过开发一种通用的数据集构建方法,并对Gemma 2模型进行后训练,旨在提升Gemma 2在一种代表性不足的语言中的性能,展示如何为本国语言和文化解锁生成式AI潜力,同时保护文化遗产。

原文摘要 · Abstract (English)

The rise of Large Language Models has not been inclusive of all cultures. The models are mostly trained on English texts and culture which makes them underperform in other languages and cultural contexts. By developing a generalizable method for preparing culturally relevant datasets and post-training the Gemma 2 model, this project aimed to increase the performance of Gemma 2 for an underrepresented language and showcase how others can do the same to unlock the power of Generative AI in their country and preserve their cultural heritage.

小语种模型微调文化适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。