用分段融合多自监督模型,提升非母语者口语流利度评估准确率。
CBF-AFA: Chunk-Based Multi-SSL Fusion for Automatic Fluency Assessment
- 按呼吸单元分段,结合多个自监督模型提取互补特征。
- 在Speechocean762上F1提升2.8,相关性提高6.2点。
- 适合需要高精度口语流利度分析的教育与语音评测场景。
自动流利度评估(AFA)在捕捉非母语者说话节奏、停顿和不流畅现象方面仍具挑战。本文提出一种基于分段的多自监督学习(SSL)融合方法,集成Wav2Vec2、HuBERT和WavLM三种模型,利用其在音位、韵律和嘈杂语音建模上的互补优势,并采用分层CNN-BiLSTM框架。通过Silero-VAD将语音划分为呼吸段,实现细粒度时序分析,同时避免过度分割伪影。SSL嵌入通过可学习加权机制融合,结合段级流利度标记(如语速、停顿时长、n-gram重复)。CNN-BiLSTM捕捉段间局部与长期依赖。在Avalinguo和Speechocean762数据集上,相比单个SSL基线,本方法在Speechocean762上F1提升2.8,皮尔逊相关性提高6.2点;在Avalinguo上分别提升4.2和4.0点,优于Pyannote.audio分段基线。结果表明,分段式多SSL融合能有效提升流利度评估鲁棒性,未来需探索方言中不规则韵律的泛化能力。
原文摘要 · Abstract (English)
Automatic fluency assessment (AFA) remains challenging, particularly in capturing speech rhythm, pauses, and disfluencies in non-native speakers. We introduce a chunk-based approach integrating self-supervised learning (SSL) models (Wav2Vec2, HuBERT, and WavLM) selected for their complementary strengths in phonetic, prosodic, and noisy speech modeling, with a hierarchical CNN-BiLSTM framework. Speech is segmented into breath-group chunks using Silero voice activity detection (Silero-VAD), enabling fine-grained temporal analysis while mitigating over-segmentation artifacts. SSL embeddings are fused via a learnable weighted mechanism, balancing acoustic and linguistic features, and enriched with chunk-level fluency markers (e.g., speech rate, pause durations, n-gram repetitions). The CNN-BiLSTM captures local and long-term dependencies across chunks. Evaluated on Avalinguo and Speechocean762, our approach improves F1-score by 2.8 and Pearson correlation by 6.2 points over single SSL baselines on Speechocean762, with gains of 4.2 F1-score and 4.0 Pearson points on Avalinguo, surpassing Pyannote.audio-based segmentation baselines. These findings highlight chunk-based multi-SSL fusion for robust fluency evaluation, though future work should explore generalization to dialects with irregular prosody.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。