arXiv:2505.24493cs.AIcs.SD2025-05被引 4

用GPT-4o仅凭文字自动标注情感语音数据,提升标注效率与一致性。

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge

  • 基于GPT-4o的结构化提示,仅用文本信息实现多模态情感标注。
  • 在Friends数据集上生成的MELT数据集使SER模型性能提升12.3%。
  • 适合需要高质量自监督情感数据的研究者,尤其关注自动化标注。

尽管深度学习推动了语音情感识别(SER)的发展,但人工标注仍是主要瓶颈。人工标注成本高且易因个体偏好或缺乏上下文知识导致标签不一致。大型语言模型(LLMs)虽可高效标注文本,但其在无监督下进行情感语音标注的潜力尚未充分探索。为此,我们利用GPT-4o仅以文本为输入,对来自情景喜剧《老友记》的多模态数据集进行标注。通过设计结构化提示,充分利用GPT-4o训练中积累的知识,证明其可在未接触多模态输入的情况下生成准确、语境相关的标注。由此提出MELT——一个完全由GPT-4o标注的多模态情感数据集。我们通过微调四个自监督学习(SSL)骨干网络,在多个情感数据集上验证了MELT的有效性,结果表明其显著提升SER性能;主观实验也证实了性能的一致性提升。

原文摘要 · Abstract (English)

Although speech emotion recognition (SER) has advanced significantly with deep learning, annotation remains a major hurdle. Human annotation is not only costly but also subject to inconsistencies annotators often have different preferences and may lack the necessary contextual knowledge, which can lead to varied and inaccurate labels. Meanwhile, Large Language Models (LLMs) have emerged as a scalable alternative for annotating text data. However, the potential of LLMs to perform emotional speech data annotation without human supervision has yet to be thoroughly investigated. To address these problems, we apply GPT-4o to annotate a multimodal dataset collected from the sitcom Friends, using only textual cues as inputs. By crafting structured text prompts, our methodology capitalizes on the knowledge GPT-4o has accumulated during its training, showcasing that it can generate accurate and contextually relevant annotations without direct access to multimodal inputs. Therefore, we propose MELT, a multimodal emotion dataset fully annotated by GPT-4o. We demonstrate the effectiveness of MELT by fine-tuning four self-supervised learning (SSL) backbones and assessing speech emotion recognition performance across emotion datasets. Additionally, our subjective experiments\' results demonstrate a consistence performance improvement on SER.

情感识别大模型应用自动化标注多模态数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。