构建多语言大学教材问答数据集,评估大模型教育应用能力
OpenStaxQA: A multilingual dataset based on open-source college textbooks
- 基于43本开源教材构建多语言问答数据集
- 70亿参数模型在该数据集上通过QLoRa微调表现良好
- 适用于教育AI研究者及多语言学习系统开发
我们提出OpenStaxQA,一个基于43本开放获取大学教材的多语言评测基准,涵盖英语、西班牙语和波兰语,均采用宽松的知识共享许可。我们在该数据集上对约70亿参数的大语言模型使用量化低秩适配器(QLoRa)进行微调与评估,并在AI2推理挑战开发集上进行零样本测试,以检验OpenStaxQA是否能提升其他任务表现。同时讨论了此类数据集带来的广泛影响。
原文摘要 · Abstract (English)
We present OpenStaxQA, an evaluation benchmark specific to college-level educational applications based on 43 open-source college textbooks in English, Spanish, and Polish, available under a permissive Creative Commons license. We finetune and evaluate large language models (LLMs) with approximately 7 billion parameters on this dataset using quantized low rank adapters (QLoRa). Additionally we also perform a zero-shot evaluation on the AI2 reasoning challenge dev dataset in order to check if OpenStaxQA can lead to an improved performance on other tasks. We also discuss broader impacts relevant to datasets such as OpenStaxQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。