arXiv:2607.07408cs.CL2026-07

用Whisper改进巴西葡萄牙语韵律边界分割,效果优于传统方法

Transformer-based segmentation of prosodic boundaries in Brazilian Portuguese

论文配图:Transformer-based segmentation of prosodic boundaries in Brazilian Portuguese
图 1 · 摘自论文原文
  • 基于Whisper大模型微调,加入显式韵律边界标记
  • 在测试集上达F1=0.731,跨数据集测试达F1=0.796
  • 能捕捉语法、语义和韵律多重线索,适合语音处理研究者

自动韵律分割通过声学与语言证据识别语音单元间的边界。尽管深度学习在英语中已取得良好效果,巴西葡萄牙语(BP)的自动分割仍主要依赖规则或传统机器学习方法。本文提出SAMPA,一个基于Whisper的分段模型,在转录BP语音时插入显式的终末韵律边界标记。我们在NURC-SP数据集的手动标注录音上微调Whisper large-v3,并评估不同训练与测试时过滤配置,包括在MuPe-Diversidades数据集上的分布外测试。SAMPA在各类设置下表现优异,最佳模型在保留测试集上达到F1=0.731,MuPe-Diversidades上达F1=0.796。通过n-gram与声学-视觉分析,我们发现模型能有效利用形态句法、语义及韵律线索进行边界检测。

原文摘要 · Abstract (English)

Automatic prosodic segmentation identifies boundaries between speech units from acoustic and linguistic evidence. Although recent deep learning approaches have produced strong results for English, automatic segmentation for Brazilian Portuguese (BP) still relies mostly on rule-based or traditional machine-learning methods. This paper presents SAMPA, a Whisper-based segmenter that transcribes BP speech while inserting explicit markers for terminal prosodic boundaries. We fine-tune Whisper large-v3 on manually segmented recordings from the NURC-SP dataset and evaluate different training and test-time filtering configurations, including out-of-distribution testing on the MuPe-Diversidades dataset. SAMPA achieves competitive boundary-detection performance across settings, with the best models reaching F1=0.731 on the held-out test split and F1=0.796 on MuPe-Diversidades. Finally, through n-gram and acoustic-visual analyses, we show that our model follows morphosyntactic, semantic, and prosodic cues for detecting prosodic boundaries.

韵律分割语音处理Whisper巴西葡语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。