构建首个统一评估框架,对比14种声学基础模型在13项非语言语音任务的表现。
ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics over Acoustic Foundation Models
- 设计统一评测流程,覆盖10个数据集的长短时非语言特征分析
- 在14个声学基础模型上测试,实现跨任务、跨模型公平比较
- 为情感计算等研究提供可复现基准,推动通用语音理解发展
计算型副语言学(ComParal)旨在自动检测、分析和解释语音中的非语言信息,如情绪、健康状态、年龄和性别。尽管进展迅速,其仍严重依赖针对特定任务精心设计的模型,导致模型异质性高,难以实际应用。近年来,随着自监督学习驱动的声学基础模型兴起,开发能高效感知多种副语言信息的通用模型成为语音处理新方向。然而,缺乏统一的评估框架制约了公平比较。为此,我们提出大规模基准ParaLBench,聚焦于不同声学基础模型在多种副语言任务上的标准化评估,涵盖情感识别与情感维度预测等关键任务。该基准包含10个数据集、13项独立任务,覆盖短、中、长期特性,所有任务均在14个声学基础模型上以统一框架执行,支持无偏方法比较,并为社区提供可靠参考。基于此,我们进一步指出跨语料泛化能力等潜在研究方向。相关代码将公开,以促进研究透明性与可复现性。
原文摘要 · Abstract (English)
Computational paralinguistics (ComParal) aims to develop algorithms and models to automatically detect, analyze, and interpret non-verbal information from speech communication, e. g., emotion, health state, age, and gender. Despite its rapid progress, it heavily depends on sophisticatedly designed models given specific paralinguistic tasks. Thus, the heterogeneity and diversity of ComParal models largely prevent the realistic implementation of ComParal models. Recently, with the advent of acoustic foundation models because of self-supervised learning, developing more generic models that can efficiently perceive a plethora of paralinguistic information has become an active topic in speech processing. However, it lacks a unified evaluation framework for a fair and consistent performance comparison. To bridge this gap, we conduct a large-scale benchmark, namely ParaLBench, which concentrates on standardizing the evaluation process of diverse paralinguistic tasks, including critical aspects of affective computing such as emotion recognition and emotion dimensions prediction, over different acoustic foundation models. This benchmark contains ten datasets with thirteen distinct paralinguistic tasks, covering short-, medium- and long-term characteristics. Each task is carried out on 14 acoustic foundation models under a unified evaluation framework, which allows for an unbiased methodological comparison and offers a grounded reference for the ComParal community. Based on the insights gained from ParaLBench, we also point out potential research directions, i.e., the cross-corpus generalizability, to propel ComParal research in the future. The code associated with this study will be available to foster the transparency and replicability of this work for succeeding researchers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。