构建多维度语音特征基准,评估说话人与语音属性的全面刻画能力
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
- 基于语音科学设计多维指标,涵盖静态(年龄、性别、口音)与动态(情绪、语流)特征
- 覆盖15+公开数据集,验证多种语音基础模型在多属性表征上的表现
- 适用于语音识别、生成系统评估及自动化评价,可与人工标注对比验证
我们提出Vox-Profile,一个用于刻画丰富说话人与语音特性的综合基准,借助语音基础模型实现。与以往仅关注单一说话人特征的研究不同,Vox-Profile提供涵盖静态特征(如年龄、性别、口音)和动态语音属性(如情绪、语流)的全维度画像。该基准基于语音科学与语言学,由领域专家共同构建,确保对说话人与语音特征的准确索引。我们使用超过15个公开语音数据集及多个主流语音基础模型进行了基准实验,涵盖多种静态与动态特征。此外,展示了多项下游应用:利用Vox-Profile增强现有语音识别数据集,分析ASR性能差异;作为工具评估语音生成系统表现;并通过与人工评估对比,验证自动画像的收敛效度。Vox-Profile已开源:https://github.com/tiantiaf0627/vox-profile-release。
原文摘要 · Abstract (English)
We introduce Vox-Profile, a comprehensive benchmark to characterize rich speaker and speech traits using speech foundation models. Unlike existing works that focus on a single dimension of speaker traits, Vox-Profile provides holistic and multi-dimensional profiles that reflect both static speaker traits (e.g., age, sex, accent) and dynamic speech properties (e.g., emotion, speech flow). This benchmark is grounded in speech science and linguistics, developed with domain experts to accurately index speaker and speech characteristics. We report benchmark experiments using over 15 publicly available speech datasets and several widely used speech foundation models that target various static and dynamic speaker and speech properties. In addition to benchmark experiments, we showcase several downstream applications supported by Vox-Profile. First, we show that Vox-Profile can augment existing speech recognition datasets to analyze ASR performance variability. Vox-Profile is also used as a tool to evaluate the performance of speech generation systems. Finally, we assess the quality of our automated profiles through comparison with human evaluation and show convergent validity. Vox-Profile is publicly available at: https://github.com/tiantiaf0627/vox-profile-release.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。