构建首个精确标注时序的语音声学地标数据集与开源工具。
Auto-Landmark: Acoustic Landmark Dataset and Open-Source Toolkit for Landmark Extraction
- 基于音素边界与人工校验标注TIMIT数据集中的声学地标。
- 首次提供精确到毫秒级的地标时序信息,支持下游任务。
- 开源提取工具与基准线,助力语音分析研究。
语音信号中的声学地标指语言特征最显著的时刻,广泛应用于语音识别、抑郁检测、临床语音异常分析及障碍性言语检测。然而现有数据集中缺乏精确的地标时间标注,而此前地标提取工具既未开源也无基准评测。本文基于前期研究选取关键声学地标,结合音素边界与人工检查对TIMIT数据集进行标注,并开发了开源的Python地标提取工具,建立了系列检测基准。该数据集、工具与基准线均为首次发布,旨在支持未来多样化研究。
原文摘要 · Abstract (English)
In the speech signal, acoustic landmarks identify times when the acoustic manifestations of the linguistically motivated distinctive features are most salient. Acoustic landmarks have been widely applied in various domains, including speech recognition, speech depression detection, clinical analysis of speech abnormalities, and the detection of disordered speech. However, there is currently no dataset available that provides precise timing information for landmarks, which has been proven to be crucial for downstream applications involving landmarks. In this paper, we selected the most useful acoustic landmarks based on previous research and annotated the TIMIT dataset with them, based on a combination of phoneme boundary information and manual inspection. Moreover, previous landmark extraction tools were not open source or benchmarked, so to address this, we developed an open source Python-based landmark extraction tool and established a series of landmark detection baselines. The first of their kinds, the dataset with landmark precise timing information, landmark extraction tool and baselines are designed to support a wide variety of future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。