arXiv:2606.23228eess.AS2026-06中稿 · Interspeech 2026 M…

用Conformer模型提升语音关键点检测准确率,解决标注差异问题。

Acoustic Landmark Detector based on Conformer and HuBERT

论文配图:Acoustic Landmark Detector based on Conformer and HuBERT
图 1 · 摘自论文原文
  • 采用高斯软标签建模标注差异,提升检测精度
  • 冻结HuBERT特征时表现最佳,20毫秒内F1达0.77
  • 塞音和擦音易检测,元音仍具挑战性

语音中的声学关键点(突变的声学变化,对应语音事件)提供了语言学上有意义的语音分析表示。本文研究基于Conformer模型的自动关键点检测,在1839个手工标注的语句上评估了14种配置,涵盖架构、损失函数、标签表示、特征提取器和数据条件,涉及八类关键点。提出每类具有10-20毫秒时间扩散的高斯软标签,相比硬标签在20毫秒内将F1值提升7.0个百分点,有效建模了标注变异性。冻结HuBERT特征在未微调情况下表现最佳(20毫秒内F1=0.77)。塞音和擦音检测可靠(F1>0.80),而元音仍具挑战(F1≈0.55)。在本数据集上,系统达到13.8%的关键点错误率(LER)。该结果与AutoLandmark(31.3%)和SpeechMark(56.5%)不直接可比,因评估数据集和指标不同。每类关键点的可检测性随事件突变程度增加,符合Stevens理论。

原文摘要 · Abstract (English)

Acoustic landmarks (abrupt acoustic changes tied to speech events) offer a linguistically grounded representation for speech analysis. We study automatic landmark detection with Conformer models, evaluating 14 configurations spanning architecture, loss, label representation, feature extractor, and data conditions on 1 839 manually annotated utterances with eight landmark types. We propose Gaussian soft labels with per-class temporal spread (sigma=10-20 ms), improving F1-at-20 ms by 7.0% absolute vs. hard labels by modeling annotation variability. Frozen HuBERT features perform best without fine-tuning (F1-at-20 ms=0.77). Stops and fricatives are reliable (F1>0.80), while vowels remain challenging (F1 approx 0.55). On our corpus, our system reaches a 13.8% Landmark Error Rate (LER). This is not directly comparable to AutoLandmark (31.3%) or SpeechMark (56.5%), evaluated on a different corpus and metric. Per-class trends show detectability increases with event abruptness, consistent with Stevens' theory.

语音分析关键点检测ConformerHuBERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。