构建首个无种子的韩语指令数据集,保障高质量与低重复。
GLAN-QnA-KR: A Seedless Taxonomy-Driven Korean Instruction Corpus
- 用无种子分类驱动方法生成,避免内容重复。
- 30万条韩语问答,近似零重复,污染检测通过率高。
- 适合研究韩语大模型训练,尤其关注数据质量者。
我们发布 GLAN-QnA-KR,一个包含 303,581 条记录的开源韩语指令问答语料库,通过无种子分类驱动的 GLAN 合成流程生成,使用 Microsoft Phi-3.5-MoE-instruct 作为生成模型(生成时间:2024-12;发布:2024-12;许可:OpenRAIL)。语料库覆盖 1,084 个英文标签的扁平学科分类,难度等级为 100–900,每条记录平均问题字符数为 313,答案字符数为 1,098。该数据集在同类规模中具备两项罕见特性:(i) 完全重复问题仅 1 条(303,581 行中),在 5,000 样本探测中,字符三元组相似度 Jaccard ≥ 0.9 的近似重复集群为零;(ii) 在 20,000 条样本上对 KMMLU、KoBEST(五个子任务)和 HAE-RAE-Bench 进行双层污染审计,最高测试-语料库问题级字符三元组 Jaccard 为 0.163,无任何测试项 Jaccard ≥ 0.7;多语言 E5 余弦相似度最高为 0.901,仅 1 项 ≥ 0.90,无 ≥ 0.95。截至发布时,据我们所知,这是 Hugging Face Hub 上可验证的最大单一管道合成韩语指令语料库,也是唯一采用无种子分类驱动协议构建的 ≥10 万行韩语语料库。本文详细记录了生成流程、语料统计、污染审计及许可边界,便于下游引用。
原文摘要 · Abstract (English)
We release GLAN-QnA-KR, a 303,581-row openly redistributable Korean instruction-QA corpus produced via the seedless taxonomy-driven GLAN synthesis pipeline with Microsoft's Phi-3.5-MoE-instruct as the producer model (generation: 2024-12; release: 2024-12; licence: OpenRAIL). The corpus spans a flat taxonomy of 1,084 English-labelled disciplines paired with Korean question/answer text, a 100-900 difficulty scale, and a median of 313 question characters and 1,098 answer characters per record. Two properties are atypical for synthetic instruction data at this scale: (i) exact duplicate questions number only 1 in 303,581 rows and character-trigram near-duplicate clusters at Jaccard >= 0.9 number zero in a 5,000-sample probe, and (ii) a two-layer contamination audit against KMMLU, KoBEST (five sub-tasks), and HAE-RAE-Bench shows a maximum test-vs-corpus question-level character-trigram Jaccard of 0.163 with zero test items at Jaccard >= 0.7, and a maximum multilingual-E5 cosine of 0.901 with a single test item at cosine >= 0.90 and zero at >= 0.95, across 20,000 sampled GLAN questions and seven evaluation sets. At the time of release, this is, to our knowledge, the largest single-pipeline synthetic Korean instruction corpus verifiable on the Hugging Face Hub and the only Korean >=100k-row corpus built under a seedless taxonomy-driven protocol. This note documents the generation protocol, corpus statistics, the contamination audit, and the licensing boundary in a form suitable for downstream citation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。