arXiv:2602.21183cs.SD2026-02

针对孟加拉语长音频,构建了语音识别与说话人分离的高效系统。

823-OLT @ BUET DL Sprint 4.0: Context-Aware Windowing for ASR and Fine-Tuned Speaker Diarization in Bengali Long Form Audio

  • 采用基于Whisper Medium的模型,结合上下文感知窗口策略提升识别稳定性。
  • 在官方竞赛数据集上微调分割模型,更好捕捉孟加拉语对话模式。
  • 适用于低资源语言的长音频转录与说话人区分,具可扩展性。

尽管孟加拉语是全球使用最广泛的语言之一,但在长音频语音技术方面仍属空白,尤其在语音转录和说话人归属任务中。本文提出面向长音频孟加拉语语音智能的框架,采用基于Whisper Medium的模型实现自动语音识别(ASR),并利用微调后的分割模型完成说话人分离。ASR流程包含声源分离、语音活动检测和间隙感知的窗口化策略,以构建保持上下文的分段,实现稳定解码。说话人分离则在由BUET CSE节主办的DL Sprint 4.0竞赛提供的官方数据集上,对预训练的说话人分割模型进行微调,更准确捕捉孟加拉语对话特征。所提系统实现了长音频的高效转录与说话人感知转录,为低资源语言提供了可扩展的语音技术解决方案。

原文摘要 · Abstract (English)

Bengali, despite being one of the most widely spoken languages globally, remains underrepresented in long form speech technology, particularly in systems addressing transcription and speaker attribution. We present frameworks for long form Bengali speech intelligence that address automatic speech recognition using a Whisper Medium based model and speaker diarization using a finetuned segmentation model. The ASR pipeline incorporates vocal separation, voice activity detection, and a gap aware windowing strategy to construct context preserving segments for stable decoding. For diarization, a pretrained speaker segmentation model is finetuned on the official competition dataset (provided as part of the DL Sprint 4.0 competition organized under BUET CSE Fest), to better capture Bengali conversational patterns. The resulting systems deliver both efficient transcription of long form audio and speaker aware transcription to provide scalable speech technology solutions for low resource languages.

语音识别说话人分离低资源语言长音频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。