arXiv:2409.05750eess.AScs.MM2024-09

一站式工具实现说话人分离与识别,支持多场景语音转写

A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR

  • 模块化设计,通过配置文件灵活集成多种模型
  • 支持多说话人、多语言及不同声学环境下的联合处理
  • 提供网页界面,实现实时音视频的说话人标注转录

我们提出一个模块化工具包,用于联合执行说话人分离与识别。该工具包可通过配置文件调用多种模型和算法,具备在多种条件下(如多个注册说话人、不同声学环境和语言)以及跨应用领域(如媒体监控、机构语音分析)运行的能力。在演示中,展示了将说话人信息与自动语音识别引擎结合,生成说话人标注转录的实际应用场景。系统通过友好的基于网页的界面,处理音频和视频输入,并根据所选配置进行处理。

原文摘要 · Abstract (English)

We present a modular toolkit to perform joint speaker diarization and speaker identification. The toolkit can leverage on multiple models and algorithms which are defined in a configuration file. Such flexibility allows our system to work properly in various conditions (e.g., multiple registered speakers' sets, acoustic conditions and languages) and across application domains (e.g. media monitoring, institutional, speech analytics). In this demonstration we show a practical use-case in which speaker-related information is used jointly with automatic speech recognition engines to generate speaker-attributed transcriptions. To achieve that, we employ a user-friendly web-based interface to process audio and video inputs with the chosen configuration.

说话人分离语音识别多说话人工具包

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。