arXiv:2603.27237cs.SDcs.AI2026-03

深度模型能从音频预测节奏感,比传统特征更准。

Can pre-trained Deep Learning models predict groove ratings?

  • 用7个先进模型提取音频嵌入,直接预测节奏感评分。
  • 不同音乐风格(放克/流行/摇滚)的节奏特征明显分离。
  • 适合做音乐信息检索和节奏感知研究的学者参考。

本研究探讨深度学习模型从音频信号直接预测节奏感及其相关感知维度的能力。我们评估了七种前沿深度学习模型在提取音频嵌入后,对节奏感评分及节奏相关问题响应的预测效果,并与传统手工特征进行对比。为进一步理解机制,方法扩展至基于音源分离的乐器分析,以分离各音乐元素的贡献。结果表明,节奏特征显著受音乐风格(放克、流行、摇滚)影响,呈现清晰区分。这说明深度音频表征能有效编码复杂且风格依赖的节奏成分,而传统特征常忽略此类细节。研究证实先进深度学习模型可捕捉多维度节奏概念,展现出表征学习在预测性音乐信息检索中的巨大潜力。

原文摘要 · Abstract (English)

This study explores the extent to which deep learning models can predict groove and its related perceptual dimensions directly from audio signals. We critically examine the effectiveness of seven state-of-the-art deep learning models in predicting groove ratings and responses to groove-related queries through the extraction of audio embeddings. Additionally, we compare these predictions with traditional handcrafted audio features. To better understand the underlying mechanics, we extend this methodology to analyze predictions based on source-separated instruments, thereby isolating the contributions of individual musical elements. Our analysis reveals a clear separation of groove characteristics driven by the underlying musical style of the tracks (funk, pop, and rock). These findings indicate that deep audio representations can successfully encode complex, style-dependent groove components that traditional features often miss. Ultimately, this work highlights the capacity of advanced deep learning models to capture the multifaceted concept of groove, demonstrating the strong potential of representation learning to advance predictive Music Information Retrieval methodologies.

节奏感深度学习音乐信息检索音频表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。