arXiv:2509.04516cs.CLcs.CY2025-09被引 1

对比英训与斯瓦希里原生训练模型,发现母语训练更准确。

Artificially Fluent: Swahili AI Performance Benchmarks Between English-Trained and Natively-Trained Datasets

  • 用斯瓦希里语原生数据训练模型,性能优于翻译成英语后用英训模型评估
  • 原生模型错误率0.36%,翻译后评估错误率达1.47%,差距近四倍
  • 提醒避免语言偏见,适合关注AI公平性与低资源语言的研究者

随着大语言模型拓展多语言能力,其在不同语言间的性能公平性仍存疑问。尽管许多群体可受益于人工智能系统,但训练数据中英语主导可能对非英语使用者造成不利影响。本研究对比了两个单语BERT模型:一个完全在斯瓦希里语数据上训练与测试,另一个在等量英语新闻数据上训练。为模拟多语言模型通过内部翻译与抽象处理非英语查询的方式,将斯瓦希里新闻数据翻译成英语,并使用英语训练模型进行评估。该方法检验了语言一致性与跨语言抽象之间的差异。结果表明,即使翻译质量高,原生斯瓦希里语训练模型的性能仍显著优于经翻译后的英语模型评估结果,错误率分别为0.36%和1.47%,差距接近四倍。这说明翻译无法弥合语言间表征差异,且英语训练模型在处理翻译输入时因内部知识表示不完整而表现不佳,凸显原生语言训练对可靠结果的重要性。在教育与信息场景中,微小性能差距也可能加剧不平等。未来研究应聚焦于欠代表语言的数据集建设与多语言模型评估,以减少全球人工智能部署对现有数字鸿沟的强化作用。

原文摘要 · Abstract (English)

As large language models (LLMs) expand multilingual capabilities, questions remain about the equity of their performance across languages. While many communities stand to benefit from AI systems, the dominance of English in training data risks disadvantaging non-English speakers. To test the hypothesis that such data disparities may affect model performance, this study compares two monolingual BERT models: one trained and tested entirely on Swahili data, and another on comparable English news data. To simulate how multilingual LLMs process non-English queries through internal translation and abstraction, we translated the Swahili news data into English and evaluated it using the English-trained model. This approach tests the hypothesis by evaluating whether translating Swahili inputs for evaluation on an English model yields better or worse performance compared to training and testing a model entirely in Swahili, thus isolating the effect of language consistency versus cross-lingual abstraction. The results prove that, despite high-quality translation, the native Swahili-trained model performed better than the Swahili-to-English translated model, producing nearly four times fewer errors: 0.36% vs. 1.47% respectively. This gap suggests that translation alone does not bridge representational differences between languages and that models trained in one language may struggle to accurately interpret translated inputs due to imperfect internal knowledge representation, suggesting that native-language training remains important for reliable outcomes. In educational and informational contexts, even small performance gaps may compound inequality. Future research should focus on addressing broader dataset development for underrepresented languages and renewed attention to multilingual model evaluation, ensuring the reinforcing effect of global AI deployment on existing digital divides is reduced.

多语言模型公平性斯瓦希里语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。