arXiv:2605.30457eess.AScs.CL2026-05

仅用声学标签提取巴葡方言特征,效果优于通用模型。

Extracting accent features in spoken Brazilian Portuguese without sociolinguistic labels

  • 基于音素强制对齐提取局部方言特征
  • 在无社会语言标签下实现更优的方言区分性能
  • 适合语音识别与方言分析研究者使用

巴西葡萄牙语(pt-BR)的区域口音分类面临可靠标注不足的问题。尽管大规模自监督学习(SSL)语音模型能力强大,但其训练流程会稀释社会语音信息,因口音标签通常不可靠或未用于训练目标。本文提出一种新工作流,仅使用声学标签进行特征提取。通过分离明确的区域口音标志,并采用基于音素的强制对齐工具(ZIPA),所生成的目标特征集能更有效地捕捉方言差异,相比语句嵌入表现更优。结果表明,在仅使用少量且客观的声学标签条件下,局部特征可超越通用架构在口音相关任务上的表现。

原文摘要 · Abstract (English)

Regional accent classification in Brazilian Portuguese (pt-BR) suffers from the need for reliable labeling. While large self-supervised learning (SSL) speech models are powerful, their training pipelines dilute sociophonetic information, since accent labels are generally not reliable or are not used in training objectives. This work introduces a novel workflow for feature extraction using only acoustic labels. By isolating explicit regional accent landmarks and using a phoneme-based forced aligner (ZIPA), our targeted feature set captures dialectal variance more effectively than utterance embeddings, demonstrating that localized features can outperform general-purpose architectures on accent-related tasks using minimal and objective data labels.

语音识别方言分析特征提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。