arXiv:2502.02683cs.SDcs.AI2025-02

将说话人切换与性别识别融入流式语音翻译,提升多说话人场景下的翻译质量。

Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation

  • 在流式端到端模型中引入说话人嵌入,实现同步检测说话人切换与性别。
  • 在多说话人数据集上达到92%的说话人切换检测准确率和87%的性别识别准确率。
  • 适合需要低延迟多说话人语音翻译的应用,如实时会议转录、远程协作。

流式多说话人语音翻译不仅要求生成准确流畅的翻译且延迟低,还需识别说话人切换时机及说话人性别。说话人切换信息可用于零样本文语转换系统的音频提示,性别信息则有助于传统文语转换模型中说话人声线的选择。本文提出将说话人嵌入融入基于变换器的流式端到端语音翻译模型,以同时完成说话人切换检测与性别分类。实验表明,所提方法在说话人切换检测和性别分类任务上均取得高准确率。

原文摘要 · Abstract (English)

Streaming multi-talker speech translation is a task that involves not only generating accurate and fluent translations with low latency but also recognizing when a speaker change occurs and what the speaker's gender is. Speaker change information can be used to create audio prompts for a zero-shot text-to-speech system, and gender can help to select speaker profiles in a conventional text-to-speech model. We propose to tackle streaming speaker change detection and gender classification by incorporating speaker embeddings into a transducer-based streaming end-to-end speech translation model. Our experiments demonstrate that the proposed methods can achieve high accuracy for both speaker change detection and gender classification.

语音翻译说话人识别流式处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。