对比中日韩语言数据集,发现三地资源建设模式差异。
No Language Data Left Behind: A Comparative Study of CJK Language Datasets in the Hugging Face Ecosystem
- 分析3300+个Hugging Face数据集,比较中日韩语言资源分布
- 中国以机构主导大规模数据,韩国靠社区自发建设,日本偏重亚文化内容
- 揭示文化与制度如何影响数据质量与共享,适合跨语言研究者参考
近年来自然语言处理的进步凸显了高质量数据集在构建大语言模型中的关键作用。然而,尽管英语资源丰富且有系统分析,东亚语言(中文、日文、韩文)的数据生态仍分散且研究不足,而这些语言合计服务超过16亿使用者。为填补这一空白,本文从跨语言视角考察Hugging Face生态系统,探究文化规范、科研环境与制度实践如何影响数据集的可获得性与质量。基于超过3,300个数据集,采用定量与定性方法分析中、日、韩三个语种社群在数据创建与管理上的不同模式。研究发现:中文数据多由机构驱动、规模庞大;韩语以社区自发为主;日语数据则高度集中于娱乐与亚文化领域。本研究揭示了提升数据文档规范性、许可证清晰度及跨语言资源共享的实用策略,为更有效、更具文化敏感性的东亚大语言模型开发提供指引。最后,提出未来数据集编纂与协作的最佳实践,旨在强化三国语言资源的整体发展。
原文摘要 · Abstract (English)
Recent advances in Natural Language Processing (NLP) have underscored the crucial role of high-quality datasets in building large language models (LLMs). However, while extensive resources and analyses exist for English, the landscape for East Asian languages - particularly Chinese, Japanese, and Korean (CJK) - remains fragmented and underexplored, despite these languages together serving over 1.6 billion speakers. To address this gap, we investigate the HuggingFace ecosystem from a cross-linguistic perspective, focusing on how cultural norms, research environments, and institutional practices shape dataset availability and quality. Drawing on more than 3,300 datasets, we employ quantitative and qualitative methods to examine how these factors drive distinct creation and curation patterns across Chinese, Japanese, and Korean NLP communities. Our findings highlight the large-scale and often institution-driven nature of Chinese datasets, grassroots community-led development in Korean NLP, and an entertainment- and subculture-focused emphasis on Japanese collections. By uncovering these patterns, we reveal practical strategies for enhancing dataset documentation, licensing clarity, and cross-lingual resource sharing - ultimately guiding more effective and culturally attuned LLM development in East Asia. We conclude by discussing best practices for future dataset curation and collaboration, aiming to strengthen resource development across all three languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。