构建10小时埃塞俄比亚圣咏数据集,助力濒危宗教音乐数字化研究
Zema Dataset: A Comprehensive Study of Yaredawi Zema with a Focus on Horologium Chants
- 构建369个实例的标注数据集,含逐字时间边界与吟诵调式标签
- 实现手稿多重记谱与实际吟诵的对应标注,解决符号-声音对齐难题
- 面向音乐信息检索、歌词转录与生成等任务,适合文化保护与音高建模研究者
计算音乐研究在推动全球各类音乐创作、传播与理解中起着关键作用。尽管埃塞俄比亚正统东正教(EOTC)圣咏具有重大的文化和宗教意义,但在计算音乐研究中仍相对匮乏。本文提出一个专为分析EOTC圣咏(又称Yaredawi Zema)设计的新数据集。该数据集包含10小时音频、369个实例,详细描述了其创建与整理过程,并实施了严格的质控措施。数据集提供逐字时间边界、诵读音调标注及对应的吟诵模式标签,同时通过标注将手稿中的多重记谱与实际吟诵进行匹配。公开此数据集旨在激励更多关于歌词转录、歌词-音频对齐及音乐生成的研究,以促进对该独特礼仪音乐的认知,守护埃塞俄比亚人民的珍贵文化遗产。
原文摘要 · Abstract (English)
Computational music research plays a critical role in advancing music production, distribution, and understanding across various musical styles worldwide. Despite the immense cultural and religious significance, the Ethiopian Orthodox Tewahedo Church (EOTC) chants are relatively underrepresented in computational music research. This paper contributes to this field by introducing a new dataset specifically tailored for analyzing EOTC chants, also known as Yaredawi Zema. This work provides a comprehensive overview of a 10-hour dataset, 369 instances, creation, and curation process, including rigorous quality assurance measures. Our dataset has a detailed word-level temporal boundary and reading tone annotation along with the corresponding chanting mode label of audios. Moreover, we have also identified the chanting options associated with multiple chanting notations in the manuscript by annotating them accordingly. Our goal in making this dataset available to the public 1 is to encourage more research and study of EOTC chants, including lyrics transcription, lyric-to-audio alignment, and music generation tasks. Such research work will advance knowledge and efforts to preserve this distinctive liturgical music, a priceless cultural artifact for the Ethiopian people.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。