提出七类讽刺分类与情感驱动生成方法,提升模型理解与创作讽刺能力。
Sarc7: Evaluating Sarcasm Detection and Generation with Seven Types and Emotion-Informed Techniques
- 基于情绪线索设计提示策略,增强讽刺识别能力
- 情感提示使模型生成成功率提升38.46%
- 首个覆盖七类讽刺的基准评测数据集
讽刺是一种表达方式,其语义与字面意义相反。利用大语言模型进行讽刺识别与生成对理解人类交流至关重要。由于讽刺具有高度微妙性,给计算模型带来挑战。我们引入Sarc7,一个通过标注MUStARD数据集构建的基准,用于分类七类讽刺:自贬型、沉思型、冷淡型、礼貌型、讨厌型、愤怒型和狂躁型。采用零样本、少样本、思维链(CoT)及一种新颖的情绪引导提示技术进行评估。我们提出一种基于情绪的生成方法,通过识别讽刺中的不一致、冲击力和上下文依赖等关键要素。分类实验表明,使用情绪提示的Gemini 2.5在零样本设置下取得0.3664的F1分数,优于其他方案。人工评估显示,情感提示生成效果比零样本多出38.46%的成功率。
原文摘要 · Abstract (English)
Sarcasm is a form of humor where expressions convey meanings opposite to their literal interpretations. Classifying and generating sarcasm using large language models is vital for interpreting human communication. Sarcasm poses challenges for computational models, due to its nuanced nature. We introduce Sarc7, a benchmark that classifies 7 types of sarcasm: self-deprecating, brooding, deadpan, polite, obnoxious, raging, and manic by annotating entries of the MUStARD dataset. Classification was evaluated using zero-shot, few-shot, chain-of-thought (CoT), and a novel emotion-based prompting technique. We propose an emotion-based generation method developed by identifying key components of sarcasm-incongruity, shock value, and context dependency. Our classification experiments show that Gemini 2.5, using emotion-based prompting, outperforms other setups with an F1 score of 0.3664. Human evaluators preferred our emotion-based prompting, with 38.46% more successful generations than zero-shot prompting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。