提出双模型方法提升幽默风格识别准确率,尤其区分亲善与攻击性幽默。
A Two-Model Approach for Humour Style Recognition
- 构建四类幽默文本数据集,含1463条标注样本。
- 双模型方法使亲善幽默识别F1提升11.61%,14种模型均获改善。
- 适用于社交媒体、文学分析等文本幽默研究场景。
幽默是人类交流的核心部分,其风格差异显著影响社交互动与心理健康。由于缺乏标准化数据集和机器学习模型,幽默风格识别面临挑战。为此,我们构建了一个新的文本数据集,涵盖1463个实例,包括四种幽默风格(自我增强、自我贬低、亲善、攻击性)及非幽默文本,长度为4至229词。研究采用多种计算方法,包括经典机器学习分类器、文本嵌入模型及DistilBERT,建立基线性能。此外,提出一种双模型方法,显著提升幽默风格识别效果,尤其在区分亲善与攻击性幽默方面表现突出。该方法在14种模型测试中实现11.61%的F1分数提升。研究成果推动了文本幽默的计算分析,为文学、社交媒体等领域的幽默研究提供了新工具。
原文摘要 · Abstract (English)
Humour, a fundamental aspect of human communication, manifests itself in various styles that significantly impact social interactions and mental health. Recognising different humour styles poses challenges due to the lack of established datasets and machine learning (ML) models. To address this gap, we present a new text dataset for humour style recognition, comprising 1463 instances across four styles (self-enhancing, self-deprecating, affiliative, and aggressive) and non-humorous text, with lengths ranging from 4 to 229 words. Our research employs various computational methods, including classic machine learning classifiers, text embedding models, and DistilBERT, to establish baseline performance. Additionally, we propose a two-model approach to enhance humour style recognition, particularly in distinguishing between affiliative and aggressive styles. Our method demonstrates an 11.61% improvement in f1-score for affiliative humour classification, with consistent improvements in the 14 models tested. Our findings contribute to the computational analysis of humour in text, offering new tools for studying humour in literature, social media, and other textual sources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。