arXiv:2606.31595cs.SDcs.DL2026-06

整合两大罗马数字乐谱数据集,实现跨标准分析对比。

Dilemmadata: On the Interoperability of Heterogeneous Roman Numeral Datasets

  • 统一使用注释级TSV格式,打通两个数据集的语法差异。
  • 合并后得1621首作品、约280万条标注,是目前最大同质化数据集。
  • 保留84首共用曲目双份分析,支持两种传统对比研究。

近年来,为支持调性和声的数据驱动研究,学界持续构建大规模罗马数字分析语料库。本文提出dilemmadata,首个将AugmentedNet数据集(AN)与远程听辨语料库(DLC)两大主流资源通过共享的注释级TSV模式实现互操作性的资源。融合过程面临四类难题:标注标准(词汇量、语法、和弦扩展、特殊功能记法)、表征方式(行定义与转换信息保留)、工具链(music21与ms3+dimcat生态不兼容)、编纂策略(曲目取舍)。通过有意识地转换、扩充与删减信息,形式化差异,保留音乐语义,并标记可能影响标注精度的变换。一致性检验与定性评估初步验证了转换后有效性,也为反思原始标准的理论假设提供基础。去重合并后,dilemmadata包含1,621首作品及约280万条注释,是当前最大同质化罗马数字语料库,虽非完美,但保留84首在两套体系中均有分析的曲目,形成可逐音比较的共享参照集。数据已发布于Zenodo,支持跨标准互操作、对比和声建模与未来编码标准优化。

原文摘要 · Abstract (English)

In recent years, there has been growing effort to annotate and collect large-scale corpora of Roman numeral analyses in support of data-driven studies in tonal harmony. We introduce dilemmadata, the first resource to reconcile two major collections, the AugmentedNet Dataset (AN) and the Distant Listening Corpus (DLC), making them interoperable through a shared note-wise TSV schema. The reconciliation confronts four families of dilemmata: annotation-standard (the two encode the same musical fact differently in terms of vocabulary size, syntax, conventions for chord extensions, inventory of special chord functions), representational (what counts as a row, and which information survives the conversion), toolchain (incompatible Python ecosystems built around music21 vs. ms3+dimcat), and curatorial (which pieces to include, exclude, or retain twice). We resolve each by deliberately transforming, augmenting, and omitting information, formalising the mismatches, preserving musical semantics, and flagging transformations that may subtly affect annotation fidelity. Consistency checks and qualitative inspections offer a preliminary assessment of post-conversion validity and a basis for critiquing the theoretical assumptions embedded in each original standard. After removing duplicates and merging the two collections, the resulting dilemmadata (1,621 pieces and aprox. 2.8 M note-wise annotations) is the largest homogeneous Roman-numeral corpus currently available, albeit far from perfect. Crucially, we retain 84 pieces common to both corpora under each of their original analyses, yielding a shared reference set in which two equally legitimate analytical traditions can be compared note-for-note over identical musical material. Released on Zenodo, dilemmadata supports interoperability, comparative harmonization modeling, and future refinement of Roman-numeral encoding standards.

音乐分析数据融合乐谱数据罗马数字

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。