通过语音特征判断网球选手胜负,准确率超90%
Sounding Like a Winner? Prosodic Differences in Post-Match Interviews
- 用语音音调、强度等特征分析选手赛后采访
- 自监督模型识别胜负准确率达90%以上
- 适合心理学、语音分析与体育科技研究者
本研究分析网球比赛后采访中胜者与败者的语音特征差异。通过提取音调、强度等声学特征,结合Wav2Vec 2.0和HuBERT等自监督学习(SSL)表示,利用机器学习分类器判断比赛结果。实验表明,基于SSL的语音表征能有效区分胜负,捕捉到与情绪状态相关的细微语音模式;同时,音调变化等传统声学线索仍是胜利的重要标志。该方法在特定数据集上实现超过90%的分类准确率,验证了语音特征在情感与状态推断中的潜力。
原文摘要 · Abstract (English)
This study examines the prosodic characteristics associated with winning and losing in post-match tennis interviews. Additionally, this research explores the potential to classify match outcomes solely based on post-match interview recordings using prosodic features and self-supervised learning (SSL) representations. By analyzing prosodic elements such as pitch and intensity, alongside SSL models like Wav2Vec 2.0 and HuBERT, the aim is to determine whether an athlete has won or lost their match. Traditional acoustic features and deep speech representations are extracted from the data, and machine learning classifiers are employed to distinguish between winning and losing players. Results indicate that SSL representations effectively differentiate between winning and losing outcomes, capturing subtle speech patterns linked to emotional states. At the same time, prosodic cues -- such as pitch variability -- remain strong indicators of victory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。