arXiv:2412.10103cs.CL2024-12被引 7

用双模数据增强提升多模态讽刺检测,效果超三模态模型。

AMuSeD: An Attentive Deep Neural Network for Multimodal Sarcasm Detection Incorporating Bi-modal Data Augmentation

  • 通过回译和定制语音合成生成新文本与音频数据。
  • 在文本-音频模态上达到81.0%的F1分数,优于三模态模型。
  • 适合研究多模态情感分析或数据稀缺场景的开发者。

有效检测讽刺需理解语境,包括语调与面部表情。然而,多模态讽刺检测因数据稀缺面临挑战。为此,我们提出AMuSeD(面向多模态讽刺检测的注意力深度神经网络),结合穆斯特尔讽刺检测数据集(MUStARD)与双模数据增强策略。第一阶段通过多种语言回译生成多样化文本;第二阶段微调基于FastSpeech 2的语音合成系统,以保留讽刺语调。辅以云端文本转语音服务,生成对应音频。我们还研究了不同注意力机制,发现自注意力在双模融合中表现最优。实验表明,该方法在文本-音频模态上实现81.0%的F1分数,超越使用三模态的模型。

原文摘要 · Abstract (English)

Detecting sarcasm effectively requires a nuanced understanding of context, including vocal tones and facial expressions. The progression towards multimodal computational methods in sarcasm detection, however, faces challenges due to the scarcity of data. To address this, we present AMuSeD (Attentive deep neural network for MUltimodal Sarcasm dEtection incorporating bi-modal Data augmentation). This approach utilizes the Multimodal Sarcasm Detection Dataset (MUStARD) and introduces a two-phase bimodal data augmentation strategy. The first phase involves generating varied text samples through Back Translation from several secondary languages. The second phase involves the refinement of a FastSpeech 2-based speech synthesis system, tailored specifically for sarcasm to retain sarcastic intonations. Alongside a cloud-based Text-to-Speech (TTS) service, this Fine-tuned FastSpeech 2 system produces corresponding audio for the text augmentations. We also investigate various attention mechanisms for effectively merging text and audio data, finding self-attention to be the most efficient for bimodal integration. Our experiments reveal that this combined augmentation and attention approach achieves a significant F1-score of 81.0% in text-audio modalities, surpassing even models that use three modalities from the MUStARD dataset.

多模态讽刺检测数据增强语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。