提升多语言场景下说话人分离与识别的鲁棒性,适配低资源环境。
The TCG CREST -- RKMVERI Submission for the NCIIPC Startup India AI Grand Challenge
- 采用多核共识谱聚类框架优化说话人分离
- 在训练数据上显著提升低资源场景下的识别准确率
- 适合需要跨语言语音分析的实际应用
本报告总结了我方团队为首届NCIIPC印度创业人工智能大奖赛设计的集成多语言语音处理流程,针对问题陈述06:语言无关的说话人识别与分离,以及后续的转录与翻译系统。核心目标是推进说话人分离技术,尤其适用于多语言及语码转换场景。重点研究了自研说话人分离(SD)系统在真实场景中的适用性,探索了鲁棒的语音活动检测(VAD)方法,并对说话人嵌入模型进行微调以提升低资源条件下的识别效果。我们采用了自研的多核共识谱聚类框架,在组织方提供的全部训练录音中显著提升了分离性能。系统还集成了说话人与语言识别、自动语音识别(ASR)和神经机器翻译模块,并通过后处理进一步增强系统鲁棒性。
原文摘要 · Abstract (English)
In this report, we summarize the integrated multilingual audio processing pipeline developed by our team for the inaugural NCIIPC Startup India AI GRAND CHALLENGE, addressing Problem Statement 06: Language-Agnostic Speaker Identification and Diarisation, and subsequent Transcription and Translation System. Our primary focus was on advancing speaker diarization, a critical component for multilingual and code-mixed scenarios. The main intent of this work was to study the real-world applicability of our in-house speaker diarization (SD) systems. To this end, we investigated a robust voice activity detection (VAD) technique and fine-tuned speaker embedding models for improved speaker identification in low-resource settings. We leveraged our own recently proposed multi-kernel consensus spectral clustering framework, which substantially improved the diarization performance across all recordings in the training corpus provided by the organizers. Complementary modules for speaker and language identification, automatic speech recognition (ASR), and neural machine translation were integrated in the pipeline. Post-processing refinements further improved system robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。