arXiv:2601.17880cs.CV2026-01被引 1

构建首个细粒度多语言多模态《古兰经》数据集,支持语音与文本深度研究。

Quran-MD: A Fine-Grained Multilingual Multimodal Dataset of the Quran

  • 按句与词级别整合阿拉伯文、英文翻译与语音,覆盖32位诵读人
  • 每词含音素对齐音频,支持发音与语义的精细分析
  • 适用于语音识别、诵读检测、个性化学习系统等场景

我们提出 Quran-MD,一个全面的《古兰经》多模态数据集,从经文和词级别融合文本、语言和音频维度。每个经文(ayah)包含原始阿拉伯文、英文翻译及音标转写。为体现《古兰经》诵读的丰富口头传统,数据集收录了32位不同诵读人的经文级音频,反映多样诵读风格与方言差异。在词级别,每个词元均配以阿拉伯文书写、英文翻译、音标转写及对齐音频,支持发音、音系与语义上下文的细粒度分析。该数据集可支撑自然语言处理、语音识别、文本转语音合成、语言学分析及数字伊斯兰研究等多种应用。通过跨诵读人连接文本与音频模态,为《古兰经》诵读与研究的计算方法提供独特资源。除支持自动语音识别、塔吉维德检测、《古兰经》文本转语音外,还为多模态嵌入、语义检索、风格迁移与个性化辅导系统奠定基础。数据集已公开于 https://huggingface.co/datasets/Buraaq/quran-audio-text-dataset。

原文摘要 · Abstract (English)

We present Quran MD, a comprehensive multimodal dataset of the Quran that integrates textual, linguistic, and audio dimensions at the verse and word levels. For each verse (ayah), the dataset provides its original Arabic text, English translation, and phonetic transliteration. To capture the rich oral tradition of Quranic recitation, we include verse-level audio from 32 distinct reciters, reflecting diverse recitation styles and dialectical nuances. At the word level, each token is paired with its corresponding Arabic script, English translation, transliteration, and an aligned audio recording, allowing fine-grained analysis of pronunciation, phonology, and semantic context. This dataset supports various applications, including natural language processing, speech recognition, text-to-speech synthesis, linguistic analysis, and digital Islamic studies. Bridging text and audio modalities across multiple reciters, this dataset provides a unique resource to advance computational approaches to Quranic recitation and study. Beyond enabling tasks such as ASR, tajweed detection, and Quranic TTS, it lays the foundation for multimodal embeddings, semantic retrieval, style transfer, and personalized tutoring systems that can support both research and community applications. The dataset is available at https://huggingface.co/datasets/Buraaq/quran-audio-text-dataset

多模态古兰经语音分析数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。