VANPY框架可自动分析语音中的性别、年龄、情绪等特征,适合声音特征研究者使用。
VANPY: Voice Analysis Framework
- 基于Python的开源语音分析框架,支持预处理到分类全流程
- 可同时识别性别、年龄、身高及情绪强度三维度(唤醒、支配、效价)
- 专为说话人特征分析设计,易扩展适配多种应用场景
语音数据在现代数字通信中日益重要,但缺乏全面的自动化语音分析工具。为此,我们开发了VANPY(Voice Analysis in Python)框架,实现语音数据的自动化预处理、特征提取与分类。该框架为开源端到端系统,专用于语音中的说话人特征表征。其设计注重可扩展性,支持新模块集成与多场景应用。目前包含超过十五个语音分析组件,涵盖音乐/语音分离、语音活动检测、说话人嵌入、声学特征提取及各类分类模型。其中四项组件由团队自研并集成:性别分类、情绪分类、年龄回归与身高回归。模型在多个数据集上表现稳健,虽未超越当前最优水平。作为概念验证,我们在电影《低俗小说》角色语音分析任务中展示了框架能力,成功提取出性别、年龄、身高、情绪类型及三维情绪强度(唤醒、支配、效价)等多重特征。
原文摘要 · Abstract (English)
Voice data is increasingly being used in modern digital communications, yet there is still a lack of comprehensive tools for automated voice analysis and characterization. To this end, we developed the VANPY (Voice Analysis in Python) framework for automated pre-processing, feature extraction, and classification of voice data. The VANPY is an open-source end-to-end comprehensive framework that was developed for the purpose of speaker characterization from voice data. The framework is designed with extensibility in mind, allowing for easy integration of new components and adaptation to various voice analysis applications. It currently incorporates over fifteen voice analysis components - including music/speech separation, voice activity detection, speaker embedding, vocal feature extraction, and various classification models. Four of the VANPY's components were developed in-house and integrated into the framework to extend its speaker characterization capabilities: gender classification, emotion classification, age regression, and height regression. The models demonstrate robust performance across various datasets, although not surpassing state-of-the-art performance. As a proof of concept, we demonstrate the framework's ability to extract speaker characteristics on a use-case challenge of analyzing character voices from the movie "Pulp Fiction." The results illustrate the framework's capability to extract multiple speaker characteristics, including gender, age, height, emotion type, and emotion intensity measured across three dimensions: arousal, dominance, and valence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。