arXiv:2409.09866cs.CLcs.AI2024-09

构建首个完整歌唱风格描述数据集,推动歌声风格自动标注研究

S2Cap: A Benchmark and a Baseline for Singing Style Captioning

  • 定义歌唱风格描述任务,构建包含多维度特征的专用数据集
  • 涵盖多种声乐、声学及人口统计特征,支持细粒度风格分析
  • 提供简单高效基线模型,适配音乐生成与音频理解场景

歌声蕴含比普通语音更丰富的信息,包括多样的发声与声学特性。然而,当前开源的歌声音频-文本数据集仅覆盖有限属性,缺乏声学特征,限制了下游任务如风格描述的应用。为填补这一空白,我们正式定义了歌唱风格描述任务,并提出S2Cap数据集,包含详细描述的歌声样本,覆盖多样化的声乐、声学及人口统计特征。基于该数据集,我们开发了一个高效且简洁的基线算法用于歌唱风格描述。数据集已公开于https://zenodo.org/records/15673764。

原文摘要 · Abstract (English)

Singing voices contain much richer information than common voices, including varied vocal and acoustic properties. However, current open-source audio-text datasets for singing voices capture only a narrow range of attributes and lack acoustic features, leading to limited utility towards downstream tasks, such as style captioning. To fill this gap, we formally define the singing style captioning task and present S2Cap, a dataset of singing voices with detailed descriptions covering diverse vocal, acoustic, and demographic characteristics. Using this dataset, we develop an efficient and straightforward baseline algorithm for singing style captioning. The dataset is available at https://zenodo.org/records/15673764.

风格描述音频数据集歌声分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。