通过音素对比分析,发现语言发音差异可被声学特征量化。
Vocal Music under Phoneme-Conditional Analysis
- 用同首歌内音素对比法,控制歌手和旋律变量。
- 九种语言中识别语音的声学差异,准确率达85.5%。
- 适合语音学、音乐信息检索研究者参考。
每种语言的声乐都具有独特的声学特质,即使无伴奏亦然。我们探究这些差异是否可测量,并与特定音素相关联。为此,提出音素条件分析方法:在相同歌曲内,对比标记音节与匹配的非标记对照,固定演唱者、旋律和音乐类型。在九种语言、数千首歌曲中,测量了五个声学维度的影响。基于这些影响构建的歌曲级声学特征,在按艺术家分组的九分类任务中,对无伴奏人声的语种识别准确率达85.5%。这种可分离性是否源于音素效应的累积尚待验证。结果表明,音系结构会在语言歌唱中留下系统且可量化的痕迹。
原文摘要 · Abstract (English)
The vocal music of each language carries a distinctive sonic identity, even without instrumental accompaniment. We ask whether these differences are measurable and traceable to specific phonemes. To tackle this question, we introduce phoneme-conditional analysis, which isolates the acoustic effect of typologically distinctive phonemes by comparing marker syllables against matched non-marker controls within the same song, holding singer, melody, and genre constant. Across nine typologically diverse languages and thousands of songs, we measure effects along five acoustic dimensions. Song-level profiles built from these effects identify the language of an unaccompanied vocal at 85.5% balanced accuracy in a nine-way classification with folds grouped by artist; whether the separability arises by accumulation of the phoneme-local effects themselves is left open. Our findings suggest that phonological structure leaves systematic and measurable traces in how each language is sung.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。