arXiv:2508.13028cs.CL2025-08被引 2

用双模讽刺检测反馈提升语音合成的讽刺表达能力

Integrating Feedback Loss from Bi-modal Sarcasm Detector for Sarcastic Speech Synthesis

  • 引入双模讽刺检测器的反馈损失优化语音合成
  • 两阶段微调使模型更精准生成讽刺语调
  • 适合人机交互与娱乐场景中的情感语音生成

讽刺语音合成对提升娱乐和人机交互中的自然性至关重要,但因讽刺语调微妙且标注数据稀缺而困难。本文提出将双模讽刺检测器的反馈损失融入文本转语音(TTS)训练,增强模型捕捉讽刺的能力。通过迁移学习,预训练于朗读语音的模型经历两阶段微调:首先在包含多种语调风格的多样化数据集上调整,其次在专注讽刺语音的数据集上进一步优化,提升生成讽刺语音的表现力。客观与主观评估均表明,该方法显著提升了合成语音的质量、自然度及讽刺感知度。

原文摘要 · Abstract (English)

Sarcastic speech synthesis, which involves generating speech that effectively conveys sarcasm, is essential for enhancing natural interactions in applications such as entertainment and human-computer interaction. However, synthesizing sarcastic speech remains a challenge due to the nuanced prosody that characterizes sarcasm, as well as the limited availability of annotated sarcastic speech data. To address these challenges, this study introduces a novel approach that integrates feedback loss from a bi-modal sarcasm detection model into the TTS training process, enhancing the model's ability to capture and convey sarcasm. In addition, by leveraging transfer learning, a speech synthesis model pre-trained on read speech undergoes a two-stage fine-tuning process. First, it is fine-tuned on a diverse dataset encompassing various speech styles, including sarcastic speech. In the second stage, the model is further refined using a dataset focused specifically on sarcastic speech, enhancing its ability to generate sarcasm-aware speech. Objective and subjective evaluations demonstrate that our proposed methods improve the quality, naturalness, and sarcasm-awareness of synthesized speech.

语音合成讽刺识别情感语音双模模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。