arXiv:2501.05989cs.CLcs.AI2025-01被引 3

用大模型修正语音翻译中的性别偏见,让女性说话者翻译更准确

Addressing speaker gender bias in large scale speech translation systems

  • 用大语言模型根据说话人性别自动修正翻译错误
  • 女性说话者翻译准确率提升70%,在MuST-SHE测试集上领先
  • 支持三种模式,适配不同场景下的性别信息需求

本研究针对大规模语音翻译系统中存在的说话人性别偏见问题,该偏见常导致不恰当或错误的翻译结果。大型语音翻译系统中普遍存在的男性偏见,通常源于基于机器翻译系统的训练数据。本文提出两步策略:首先,利用大语言模型以低成本方式根据说话人性别纠正翻译内容;其次,使用修正后的数据微调语音翻译模型,使模型能直接从音频线索生成性别相关翻译,无需显式输入性别信息。此外,我们还设计了三种模式的微调模型,适用于说话人性别已知或不应从语音线索推断的场景。在MuST-SHE测试集上,相比基线模型及其他大型语音翻译系统(如Seamless M4T和Canary),女性说话者的翻译质量提升达70%。

原文摘要 · Abstract (English)

This study addresses the issue of speaker gender bias in Speech Translation (ST) systems, which can lead to offensive and inaccurate translations. The masculine bias often found in large-scale ST systems is typically perpetuated through training data derived from Machine Translation (MT) systems. Our approach involves two key steps. First, we employ Large Language Models (LLMs) to rectify translations based on the speaker's gender in a cost-effective manner. Second, we fine-tune the ST model with the corrected data, enabling the model to generate gender-specific translations directly from audio cues, without the need for explicit gender input. Additionally, we propose a three-mode fine-tuned model for scenarios where the speaker's gender is either predefined or should not be inferred from speech cues. We demonstrate a 70% improvement in translations for female speakers compared to our baseline and other large-scale ST systems, such as Seamless M4T and Canary, on the MuST-SHE test set.

语音翻译性别偏见大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。