用物理模型合成英语发音,构建了超20万词的声学地标数据库。
An Acoustic Landmark Database of the English Lexicon via Articulatory Synthesis

- 通过物理声腔模型从发音动作反推语音,精准标注声学地标时间点。
- 生成超过20万词的英语语音数据,涵盖男女声型,可测语音可懂度。
- 提供发音事件频率统计,适合语音分析与自动地标检测研究者使用。
声学地标理论认为语音由塑造声道和气流的发音动作所决定。但因缺乏大规模、无歧义标注的地标数据集,研究进展受限。本文通过反向建模,从地标模式生成语音。利用Pink Trombone物理声道合成器,生成两个成年发音配置(男、女)的英语词库。通过直接控制发音动作,在其真实发生时刻算法化标注地标标签(如口腔闭合/释放)。语料库包含超过20万条合成单词,针对两种发音配置均有时间对齐标注,语音可懂度通过STOI指标衡量。该数据集支持从发音事件视角进行全词库统计分析,报告地标出现频率与主导线索模式,并可用于自动地标检测器的训练与基准测试。
原文摘要 · Abstract (English)
Acoustic landmark theory treats speech as organized around the acoustic consequences of articulatory gestures that shape the vocal tract and airflow. Progress is limited by the scarcity of large, unambiguously annotated landmark datasets. We invert the problem by generating speech from landmark patterns. Using the Pink Trombone physical vocal-tract synthesizer, we produce an English lexicon for two adult configurations (male, female). With direct control of gestures, we place landmark labels algorithmically at the exact times of their physical events (e.g., oral closures/releases). The corpus contains $>$200,000 synthesized words, rendered for both configurations with time-aligned annotations; intelligibility is measured with STOI. We leverage it for statistics across the lexicon from an articulatory-event view, reporting landmark frequencies and dominant cue patterns, and enabling quantitative studies plus training/benchmarking of automatic landmark detectors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。