用视频自动生成高一致性情感语音标签,效率远超人工。
MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling
- 基于多模态大模型自动分析视频中的情绪。
- 在MELD数据集上达68.5%准确率,一致性评分为0.93。
- 适合做细粒度情感语音合成与音色克隆的研究者使用。
获取大规模、高一致性的情感语音数据仍是语音合成的挑战。本文提出MIKU-PAL,一种全自动多模态管道,可从无标注视频中提取高质量情感语音。利用人脸检测与跟踪算法,结合多模态大语言模型(MLLM),构建了自动情绪分析系统。实验表明,MIKU-PAL在MELD数据集上达到68.5%的人类级准确率,一致性评分高达0.93 Fleiss kappa,且成本与速度远优于人工标注。基于该系统,可标注多达26种细粒度语音情绪类别,经人工验证理性度达83%。进一步发布了包含131.2小时音频的细粒度情感语音数据集MIKU-EmoBench,作为情感文语转换与视觉声线克隆的新基准。
原文摘要 · Abstract (English)
Acquiring large-scale emotional speech data with strong consistency remains a challenge for speech synthesis. This paper presents MIKU-PAL, a fully automated multimodal pipeline for extracting high-consistency emotional speech from unlabeled video data. Leveraging face detection and tracking algorithms, we developed an automatic emotion analysis system using a multimodal large language model (MLLM). Our results demonstrate that MIKU-PAL can achieve human-level accuracy (68.5% on MELD) and superior consistency (0.93 Fleiss kappa score) while being much cheaper and faster than human annotation. With the high-quality, flexible, and consistent annotation from MIKU-PAL, we can annotate fine-grained speech emotion categories of up to 26 types, validated by human annotators with 83% rationality ratings. Based on our proposed system, we further released a fine-grained emotional speech dataset MIKU-EmoBench(131.2 hours) as a new benchmark for emotional text-to-speech and visual voice cloning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。