arXiv:2411.00195cs.SDcs.LG2024-11被引 8

用音频特征预测用户情感评分,可评估音乐翻唱质量。

Machine Learning Framework for Audio-Based Content Evaluation using MFCC, Chroma, Spectral Contrast, and Temporal Feature Engineering

  • 提取MFCC、Chroma等四类音频特征,分段处理30秒音频
  • 回归模型预测情感分,最低误差仅2.783(RMSE)
  • 适合音乐内容分析与AI媒体评估场景

本研究提出一种基于机器学习的音频内容评估框架,用于衡量音频相似性并预测情感得分。构建包含YouTube音乐翻唱及其原曲音频的数据集,并通过用户评论提取情感分数作为内容质量的代理标签。对音频信号进行预处理,划分为30秒窗口,提取梅尔频率倒谱系数(MFCC)、Chroma、频谱对比度和时序特征等高维特征表示。利用这些特征训练回归模型,在0-100分尺度上预测情感得分,取得3.420、5.482、2.783和4.212的均方根误差(RMSE),优于基于绝对差值的基线模型。结果表明,机器学习能有效捕捉音频中的情感与相似性信息,为媒体分析中的AI应用提供可扩展框架。

原文摘要 · Abstract (English)

This study presents a machine learning framework for assessing similarity between audio content and predicting sentiment score. We construct a dataset containing audio samples from music covers on YouTube along with the audio of the original song, and sentiment scores derived from user comments, serving as proxy labels for content quality. Our approach involves extensive pre-processing, segmenting audio signals into 30-second windows, and extracting high-dimensional feature representations through Mel-Frequency Cepstral Coefficients (MFCC), Chroma, Spectral Contrast, and Temporal characteristics. Leveraging these features, we train regression models to predict sentiment scores on a 0-100 scale, achieving root mean square error (RMSE) values of 3.420, 5.482, 2.783, and 4.212, respectively. Improvements over a baseline model based on absolute difference metrics are observed. These results demonstrate the potential of machine learning to capture sentiment and similarity in audio, offering an adaptable framework for AI applications in media analysis.

音频分析情感预测特征工程机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。