arXiv:2506.19335cs.SD2025-06中稿 · on EUSIPCO 2024被引 1

用对比评分训练模型评估语音主观印象,效果优于传统方法。

Learning to assess subjective impressions from speech

  • 基于对比评分构建个性化语音印象评估框架。
  • 对比评分训练下模型性能在少量数据时仍达中等以上水平。
  • 适合语音情感分析、个性化推荐等场景使用。

我们提出一项新任务:训练神经网络模型以评估语音传达的主观印象并打分,灵感来自自动语音质量评估(SQA)。语音印象常以‘可爱的声音’等短语描述,我们称其为主观语音描述符(SVDs)。针对该任务与SQA在使用场景上的差异,设计了支持个体化SVD(如‘我最喜欢的声音’)的框架。本研究构建了一个包含绝对类别评分(ACR)和比较类别评分(CCR)的语音标签数据集。评估指标采用ppref,即在CCR测试样本上预测得分排序的准确率。除传统基于ACR的模型与学习方法外,还探索了基于CCR的RankNet学习。实验发现,即使训练数据极少,ppref也保持中等水平;且CCR训练显著优于ACR训练。结果表明,基于个性化SVD的评估模型,虽通常面临数据有限问题,但仍可有效通过CCR数据学习。

原文摘要 · Abstract (English)

We tackle a new task of training neural network models that can assess subjective impressions conveyed through speech and assign scores accordingly, inspired by the work on automatic speech quality assessment (SQA). Speech impressions are often described using phrases like `cute voice.' We define such phrases as subjective voice descriptors (SVDs). Focusing on the difference in usage scenarios between the proposed task and automatic SQA, we design a framework capable of accommodating SVDs personalized to each individual, such as `my favorite voice.' In this work, we compiled a dataset containing speech labels derived from both abosolute category ratings (ACR) and comparison category ratings (CCR). As an evaluation metric for assessment performance, we introduce ppref, the accuracy of the predicted score ordering of two samples on CCR test samples. Alongside the conventional model and learning methods based on ACR data, we also investigated RankNet learning using CCR data. We experimentally find that the ppref is moderate even with very limited training data. We also discover the CCR training is superior to the ACR training. These results support the idea that assessment models based on personalized SVDs, which typically must be trained on limited data, can be effectively learned from CCR data.

语音评估主观印象对比学习个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。