用对比学习提升韩语声调分类准确率,解决真实语音中的音高波动问题。
Deep Supervised Contrastive Learning of Pitch Contours for Robust Pitch Accent Classification in Seoul Korean

- 通过双视角对比学习,捕捉音高轮廓的整体形状特征。
- 在10,093个标注语料上达到77.75%准确率和51.54%F1分数。
- 首个大规模韩语声调标注数据集,适合语音分析与韵律建模研究者。
首尔韩语的语调结构基于自段落-音高模型中的离散音调类别定义。然而,由于真实语音中基频(F₀)实现的可变性,将连续的F₀轮廓映射到这些不变类别极具挑战。本文提出Dual-Glob,一种深度监督对比学习框架,用于鲁棒地分类首尔韩语的细粒度声调模式。不同于传统局部预测模型,该方法通过在共享潜在空间中强制干净与增强视图之间的结构一致性,捕捉完整的F₀轮廓形态。为此,我们构建了首个大规模基准数据集,包含10,093个手动标注的声调短语。实验表明,Dual-Glob显著优于强基线模型,在准确率(77.75%)和F1分数(51.54%)上达到当前最优水平。因此,本工作以数据驱动方式支持基于AM的语调音系学,证明深度对比学习能有效捕捉连续F₀轮廓的全局结构特征。
原文摘要 · Abstract (English)
The intonational structure of Seoul Korean has been defined with discrete tonal categories within the Autosegmental-Metrical model of intonational phonology. However, it is challenging to map continuous $F_0$ contours to these invariant categories due to variable $F_0$ realizations in real-world speech. Our paper proposes Dual-Glob, a deep supervised contrastive learning framework to robustly classify fine-grained pitch accent patterns in Seoul Korean. Unlike conventional local predictive models, our approach captures holistic $F_0$ contour shapes by enforcing structural consistency between clean and augmented views in a shared latent space. To this aim, we introduce the first large-scale benchmark dataset, consisting of manually annotated 10,093 Accentual Phrases in Seoul Korean. Experimental results show that our Dual-Glob significantly outperforms strong baseline models with state-of-the-art accuracy (77.75%) and F1-score (51.54%). Therefore, our work supports AM-based intonational phonology using data-driven methodology, showing that deep contrastive learning effectively captures holistic structural features of continuous $F_0$ contours.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。