arXiv:2412.20057cs.CL2024-12ACL被引 2

识别伪装成抱怨的自夸,让机器读懂话里有话

"My life is miserable, have to sign 500 autographs everyday": Exposing Humblebragging, the Brags in Disguise

  • 用四元组定义自夸伪装现象,构建可计算检测框架
  • 人类准确率仅64%,最佳模型F1达0.88,仍具挑战性
  • 发布3340条数据集HB-24,推动自然语言理解发展

自夸伪装(Humblebragging)是人们以谦逊或抱怨形式表达自我宣传的现象。例如‘唉,我居然升职带整个团队,压力好大!’表面抱怨实则炫耀成就。该现象对情感分析与意图识别等任务至关重要,但尚未被计算语言学系统研究。本文首次提出自动检测文本中自夸伪装的任务,通过四元组定义该现象,并评估机器学习、深度学习及大语言模型(LLMs)的表现,同时与人类对比。我们构建并公开了名为HB-24的数据集,包含3340条由GPT-4o生成的自夸伪装语句。实验表明,即使对人类也难以准确识别,最佳模型的F1分数达到0.88。本工作为深入理解这一语言现象奠定基础,并推动其在更广泛自然语言理解系统中的应用。

原文摘要 · Abstract (English)

Humblebragging is a phenomenon in which individuals present self-promotional statements under the guise of modesty or complaints. For example, a statement like, "Ugh, I can't believe I got promoted to lead the entire team. So stressful!", subtly highlights an achievement while pretending to be complaining. Detecting humblebragging is important for machines to better understand the nuances of human language, especially in tasks like sentiment analysis and intent recognition. However, this topic has not yet been studied in computational linguistics. For the first time, we introduce the task of automatically detecting humblebragging in text. We formalize the task by proposing a 4-tuple definition of humblebragging and evaluate machine learning, deep learning, and large language models (LLMs) on this task, comparing their performance with humans. We also create and release a dataset called HB-24, containing 3,340 humblebrags generated using GPT-4o. Our experiments show that detecting humblebragging is non-trivial, even for humans. Our best model achieves an F1-score of 0.88. This work lays the foundation for further exploration of this nuanced linguistic phenomenon and its integration into broader natural language understanding systems.

语言理解自夸伪装情感分析大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。