用25个信号指标构建24种语音质量的客观量化公式。
Objective Measurements of Voice Quality
- 基于文献提炼25个信号特征与24种语音子质量的数学关系
- 在主观标注数据集上验证公式有效性,可精准映射感知质量
- 适合语音分析、临床评估及跨领域语音质量研究者
语音质量在音乐、言语治疗和通信等领域至关重要,但缺乏统一的客观定义,通常依赖'沙哑''气声'等主观描述。尽管如此,大量跨学科研究已将这些语音特征与说话人健康、生理特征等信息相关联。当前机器学习方法在语音画像中多采用数据驱动分析,未能充分融合这些已有定性关联。本文旨在通过整合文献中已确立的语音质量与信号处理指标间的对应关系,建立公式化表示,提出24种语音子质量的量化公式,基于25个信号属性。这些公式在带有主观标签的语音数据集上进行了验证,结果表明其具备良好有效性。
原文摘要 · Abstract (English)
The quality of human voice plays an important role across various fields like music, speech therapy, and communication, yet it lacks a universally accepted, objective definition. Instead, voice quality is referred to using subjective descriptors like "rough", "breathy" etc. Despite this subjectivity, extensive research across disciplines has linked these voice qualities to specific information about the speaker, such as health, physiological traits, and others. Current machine learning approaches for voice profiling rely on data-driven analysis without fully incorporating these established correlations, due to their qualitative nature. This paper aims to objectively quantify voice quality by synthesizing formulaic representations from past findings that correlate voice qualities to signal-processing metrics. We introduce formulae for 24 voice sub-qualities based on 25 signal properties, grounded in scientific literature. These formulae are tested against datasets with subjectively labeled voice qualities, demonstrating their validity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。