分析非洲低资源语言语料库的版权兼容性问题,揭示多个数据集存在法律隐患。
Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages
- 构建六级许可证兼容矩阵,评估语料库许可合规性。
- 发现4类问题:禁止使用、许可误导、隐藏条款、数据链接失效。
- 适合关注非洲语言数据伦理与法律合规的研究者参考。
创意共享许可证主导非洲自然语言处理语料库发布,但其兼容规则极少被遵循。CC-BY-SA与CC-BY-NC无法共存于同一数据集;NoDerivs条款隐性禁止分词与标注。本文审计了二十多个用于非洲NLP的语料库家族,构建六级兼容性矩阵,并以基图巴/穆努库图巴、扎尔马、莫罗为例进行案例研究。记录四类失败模式:直接禁止(JW300因违反服务条款被从OPUS移除);复合许可误导(WAXAL声称采用CC-BY 4.0,但HuggingFace数据卡内容矛盾);在CC-BY标签下隐藏NoDerivs条款(Tanzil);数据持久性失败(刚果广播语料库中402/405个源链接已失效)。论文末尾提供预标注尽职调查清单及合法数据增补机会调研。
原文摘要 · Abstract (English)
Creative Commons licenses dominate African NLP corpus releases, but their compatibility rules are rarely applied. CC-BY-SA and CC-BY-NC cannot be combined in a single published dataset; a NoDerivs clause silently prohibits tokenisation and annotation. This paper audits the license provenance of over twenty corpus families used in African NLP, constructs a six-tier compatibility matrix, and applies it to three case-study languages: Kituba/Munukutuba, Zarma, and Moore. Four failure modes are documented with primary-source evidence: outright prohibition (JW300, removed from OPUS after a legal audit confirmed Terms of Service violation); composite license misrepresentation (WAXAL, whose CC-BY 4.0 claim is contradicted by its own HuggingFace dataset card); a NoDerivs clause hidden behind a CC-BY label (Tanzil); and data persistence failure (the Congolese Radio Corpus, where 402 of 405 source URLs are now dead). A pre-annotation due diligence checklist and a survey of legally clean enrichment opportunities close the paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。