构建首个覆盖阿拉伯多方言的大型歌词诗集数据集,支持跨语言与历史分析。
Tarab: A Multi-Dialect Corpus of Arabic Lyrics and Poetry
- 整合256万行歌词与诗歌,涵盖古典到现代多种阿拉伯语变体
- 包含1350万以上词元,覆盖14个世纪及28国创作背景
- 适合研究阿拉伯语文学、方言差异与跨时代文化演变的学者
我们提出Tarab语料库,一个大规模的文化与语言资源,将阿拉伯歌曲歌词与诗歌统一在分析框架中。该语料库包含256万行诗句和超过1350万词元,据我们所知是目前最大的开放阿拉伯创意文本语料库,涵盖古典与当代创作。语料库在歌曲与诗歌之间保持广泛平衡,覆盖古典阿拉伯语、现代标准阿拉伯语(MSA)以及六种主要地区变体:埃及、海湾、黎凡特、伊拉克、苏丹和马格里布阿拉伯语。创作者来自28个现代国家及多个历史时期,涵盖从伊斯兰前时期到21世纪的1400多年阿拉伯创造性表达。每行诗句均附有结构化元数据,包括语言变体、地理来源和历史/文化背景,支持跨文体、跨时间的比较分析。我们描述了数据收集、标准化与验证流程,并提供了变体识别与体裁区分的基准分析。数据集已公开发布于HuggingFace:https://huggingface.co/datasets/drelhaj/Tarab。
原文摘要 · Abstract (English)
We introduce the Tarab Corpus, a large-scale cultural and linguistic resource that brings together Arabic song lyrics and poetry within a unified analytical framework. The corpus comprises 2.56 million verses and more than 13.5 million tokens, making it, to our knowledge, the largest open Arabic corpus of creative text spanning both classical and contemporary production. Tarab is broadly balanced between songs and poems and covers Classical Arabic, Modern Standard Arabic (MSA), and six major regional varieties: Egyptian, Gulf, Levantine, Iraqi, Sudanese, and Maghrebi Arabic. The artists and poets represented in the corpus are associated with 28 modern nation states and multiple historical eras, covering over fourteen centuries of Arabic creative expression from the Pre-Islamic period to the twenty-first century. Each verse is accompanied by structured metadata describing linguistic variety, geographic origin, and historical or cultural context, enabling comparative linguistic, stylistic, and diachronic analysis across genres and time. We describe the data collection, normalisation, and validation pipeline and present baseline analyses for variety identification and genre differentiation. The dataset is publicly available on HuggingFace at https://huggingface.co/datasets/drelhaj/Tarab.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。