arXiv:2507.14638cs.SDeess.AS2025-07

用生态学模型估算音乐文献中未被发现的作曲家和乐谱数量。

The Rest is Silence: Leveraging Unseen Species Models for Computational Musicology

  • 引入生态学中的未见物种模型分析音乐数据
  • 揭示了如RISM数据库中缺失的作曲家比例
  • 适合音乐信息检索与数字音乐学研究者

长期以来,音乐学家建立了大量服务于研究与学术的数据库。随着音乐信息检索和数字音乐学的发展,相关数据集与语料库持续增长。然而,在历史或观察性研究中,这些数据集必然不完整,其真实规模仍处于“沉默”状态。本文首次将生态学中的未见物种模型(Unseen Species Models, USMs)应用于音乐学领域。通过四个案例研究,展示了如何利用USMs解决量化问题:例如,RISM中尚有多少作曲家未被收录?格里高利圣咏的中世纪乐谱中有多少已编目?不同版本印刷乐谱间预期存在多少差异?民间音乐传统中歌曲覆盖程度如何?以及对大量作曲家和声词汇量的估计接近完成度?

原文摘要 · Abstract (English)

For many decades, musicologists have engaged in creating large databases serving different purposes for musicological research and scholarship. With the rise of fields like music information retrieval and digital musicology, there is now a constant and growing influx of musicologically relevant datasets and corpora. In historical or observational settings, however, these datasets are necessarily incomplete, and the true extent of a collection of interest remains unknown -- silent. Here, we apply, for the first time, so-called Unseen Species models (USMs) from ecology to areas of musicological activity. After introducing the models formally, we show in four case studies how USMs can be applied to musicological data to address quantitative questions like: How many composers are we missing in RISM? What percentage of medieval sources of Gregorian chant have we already cataloged? How many differences in music prints do we expect to find between editions? How large is the coverage of songs from genres of a folk music tradition? And, finally, how close are we in estimating the size of the harmonic vocabulary of a large number of composers?

音乐信息检索数字音乐学统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。