arXiv:2410.18882cs.CL2024-10IJCAI综述被引 44

首篇系统综述多模态讽刺检测,涵盖2018-2023年主流方法与数据集。

A Survey of Multimodal Sarcasm Detection

  • 梳理文本、语音、图像等多模态信息融合的讽刺检测方法
  • 总结近年常用数据集与模型性能,揭示跨模态依赖关系
  • 适合从事情感分析、人机交互与社交媒体研究者参考

讽刺是一种修辞手法,用于表达话语字面意义的反义。它广泛出现在社交媒体及其他计算机中介交流中,推动了自动识别讽刺的计算模型发展。尽管现有大部分方法仅基于文本,但讽刺识别常需语调、面部表情和上下文图像等额外信息。这促使多模态模型兴起,使音频、图像、文本、视频等多种模态的讽刺检测成为可能。本文首次系统综述多模态讽刺检测(MSD),覆盖2018至2023年间相关论文,分析所用模型与数据集,并提出未来研究方向。

原文摘要 · Abstract (English)

Sarcasm is a rhetorical device that is used to convey the opposite of the literal meaning of an utterance. Sarcasm is widely used on social media and other forms of computer-mediated communication motivating the use of computational models to identify it automatically. While the clear majority of approaches to sarcasm detection have been carried out on text only, sarcasm detection often requires additional information present in tonality, facial expression, and contextual images. This has led to the introduction of multimodal models, opening the possibility to detect sarcasm in multiple modalities such as audio, images, text, and video. In this paper, we present the first comprehensive survey on multimodal sarcasm detection - henceforth MSD - to date. We survey papers published between 2018 and 2023 on the topic, and discuss the models and datasets used for this task. We also present future research directions in MSD.

多模态讽刺检测综述情感分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。