构建超1400小时多风格古兰经诵读音频数据集
Tadabur: A Large-Scale Quran Audio Dataset

- 收集600多位诵读人超过1400小时古兰经音频
- 涵盖多样诵读风格、嗓音特征与录制条件
- 适合语音识别与宗教文本研究者使用
尽管古兰经数据研究日益受到关注,现有数据集在规模和多样性上仍显不足。为此,我们提出Tadabur,一个大规模古兰经音频数据集。Tadabur包含超过600位不同诵读人提供的1400+小时诵读音频,涵盖丰富的诵读风格、嗓音特征和录制条件差异。该数据集为古兰经语音研究与分析提供了全面且具有代表性的资源。通过显著扩充可用古兰经数据的总时长与多样性,Tadabur旨在支持未来研究,并推动标准化古兰经语音基准的建立。
原文摘要 · Abstract (English)
Despite growing interest in Quranic data research, existing Quran datasets remain limited in both scale and diversity. To address this gap, we present Tadabur, a large-scale Quran audio dataset. Tadabur comprises more than 1400+ hours of recitation audio from over 600 distinct reciters, providing substantial variation in recitation styles, vocal characteristics, and recording conditions. This diversity makes Tadabur a comprehensive and representative resource for Quranic speech research and analysis. By significantly expanding both the total duration and variability of available Quran data, Tadabur aims to support future research and facilitate the development of standardized Quranic speech benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。