arXiv:2410.12668cs.SDeess.AS2024-10中稿 · IEEE SLT 2024被引 3

为语音识别常用数据集VoxCeleb添加1251名说话人身高信息,助力语音身高预测研究。

HeightCeleb - an enrichment of VoxCeleb dataset with speaker height information

  • 基于公开资料自动标注VoxCeleb中1251名说话人身高,构建HeightCeleb数据集。
  • 仅用统计回归+预训练声纹嵌入,在TIMIT上达到当前最佳效果。
  • 无需微调即可利用现有模型,适合语音取证与说话人画像研究者。

说话人身高预测在语音司法鉴定、监控和自动说话人画像中有重要意义。以往研究多依赖于TIMIT数据集进行训练与评估。本文提出HeightCeleb,作为对常用说话人识别数据集VoxCeleb的扩展,新增了1251名说话人的身高信息,这些信息通过自动化方法从公开来源提取。该标注数据使研究者可直接使用在VoxCeleb上预训练的声纹嵌入提取器,构建更高效的身高估计算法。本文详细描述了HeightCeleb的构建过程,并证明仅使用简单统计回归方法与主流说话人模型(未进行额外微调)即可在TIMIT测试集上取得当前最优结果。

原文摘要 · Abstract (English)

Prediction of speaker's height is of interest for voice forensics, surveillance, and automatic speaker profiling. Until now, TIMIT has been the most popular dataset for training and evaluation of the height estimation methods. In this paper, we introduce HeightCeleb, an extension to VoxCeleb, which is the dataset commonly used in speaker recognition tasks. This enrichment consists in adding information about the height of all 1251 speakers from VoxCeleb that has been extracted with an automated method from publicly available sources. Such annotated data will enable the research community to utilize freely available speaker embedding extractors, pre-trained on VoxCeleb, to build more efficient speaker height estimators. In this work, we describe the creation of the HeightCeleb dataset and show that using it enables to achieve state-of-the-art results on the TIMIT test set by using simple statistical regression methods and embeddings obtained with a popular speaker model (without any additional fine-tuning).

语音分析数据集身高估计说话人识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。