用LSTM自动为波斯语歌词生成旋律,效果接近人工创作。
Vocal Melody Construction for Persian Lyrics Using LSTM Recurrent Neural Networks
- 基于音节与旋律的语音相关性,构建序列到序列模型。
- 模型生成旋律平均得分3.005(满分5),接近人工4.078。
- 适合音乐生成、跨语言旋律合成研究者参考。
本文研究了以波斯语歌词为输入的自动旋律生成问题,假设歌词音节与旋律存在语音关联。构建了一个序列到序列神经网络,在大量波斯歌曲的音节与音符序列对上进行训练,以生成新歌词对应的悦耳旋律。由于缺乏波斯数字音乐数据集,研究收集并数字化了100多首波斯歌曲。最后,将14组新歌词输入模型,由音乐专家演奏并录制生成旋律,通过包含170多名受试者的音频问卷评估。结果显示,系统输出平均愉悦度得分为3.005(满分5),而人类创作的同歌词旋律平均得分为4.078。
原文摘要 · Abstract (English)
The present paper investigated automatic melody construction for Persian lyrics as an input. It was assumed that there is a phonological correlation between the lyric syllables and the melody in a song. A seq2seq neural network was developed to investigate this assumption, trained on parallel syllable and note sequences in Persian songs to suggest a pleasant melody for a new sequence of syllables. More than 100 pieces of Persian music were collected and converted from the printed version to the digital format due to the lack of a dataset on Persian digital music. Finally, 14 new lyrics were given to the model as input, and the suggested melodies were performed and recorded by music experts to evaluate the trained model. The evaluation was conducted using an audio questionnaire, which more than 170 persons answered. According to the answers about the pleasantness of melody, the system outputs scored an average of 3.005 from 5, while the human-made melodies for the same lyrics obtained an average score of 4.078.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。