构建首个不丹语-英语双语问答数据集,测试大模型在低资源语言的表现
Multilingual Question Answering in Low-Resource Settings: A Dzongkha-English Benchmark for Foundation Models
- 创建覆盖5000+题的不丹语与英语平行问答数据集
- 发现大模型在不丹语上的表现显著低于英语,尤其在事实类问题上
- 英文翻译能提升不丹语回答的准确率,链式思维提示对推理题有效
本文构建了名为DZEN的平行双语问答数据集,包含超过5000个面向不丹中高学生、涵盖多种科学主题的不丹语与英语测试题,包含事实型、应用型和推理型题目。我们利用该数据集评估多个大语言模型(LLMs)在两种语言下的表现,发现模型在英语中的性能远高于不丹语。研究还表明,链式思维(CoT)提示对推理类问题效果较好,但对事实类问题帮助有限;同时,加入英文翻译可显著提升不丹语问题的回答精度。结果揭示了改进低资源语言大模型性能的新方向。数据集已公开于:https://github.com/kraritt/llm_dzongkha_evaluation。
原文摘要 · Abstract (English)
In this work, we provide DZEN, a dataset of parallel Dzongkha and English test questions for Bhutanese middle and high school students. The over 5K questions in our collection span a variety of scientific topics and include factual, application, and reasoning-based questions. We use our parallel dataset to test a number of Large Language Models (LLMs) and find a significant performance difference between the models in English and Dzongkha. We also look at different prompting strategies and discover that Chain-of-Thought (CoT) prompting works well for reasoning questions but less well for factual ones. We also find that adding English translations enhances the precision of Dzongkha question responses. Our results point to exciting avenues for further study to improve LLM performance in Dzongkha and, more generally, in low-resource languages. We release the dataset at: https://github.com/kraritt/llm_dzongkha_evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。