arXiv:2409.10760cs.CLcs.SI2024-09被引 4

用语义一致性评估新框架,让表情包推荐更贴近真实社交行为

Semantics Preserving Emoji Recommendation with Large Language Models

  • 提出语义保真评估框架,判断推荐表情是否保留原文情感与立场
  • GPT-4o表现最佳,语义保真得分达79.23%
  • 揭示模型在下游分类中的偏差,评估推荐多样性

表情符号已成为数字交流的重要组成部分,通过传递情绪、语气和意图丰富文本。现有表情包推荐方法主要依据是否匹配用户原始文本中选择的精确表情符号进行评估,但忽略了社交媒体上一个文本可对应多个合理表情的真实行为模式。为更准确衡量模型在真实场景下的表现,本文提出一种新的语义保真评估框架,衡量模型推荐的表情是否保持与原文一致的语义特征,包括预测的情感状态、人口统计特征和态度立场。若这些属性不变,则认为语义被保留。大型语言模型(LLMs)在理解与生成上下文相关、细微表达方面的优势使其非常适合处理这一任务。为此,我们构建了一个综合基准,系统评估六种专有及开源大模型在不同提示策略下的表现。实验表明,GPT-4o优于其他模型,语义保真得分为79.23%。此外,还通过案例研究分析了模型在下游分类任务中的偏差,并评估了推荐表情的多样性。

原文摘要 · Abstract (English)

Emojis have become an integral part of digital communication, enriching text by conveying emotions, tone, and intent. Existing emoji recommendation methods are primarily evaluated based on their ability to match the exact emoji a user chooses in the original text. However, they ignore the essence of users' behavior on social media in that each text can correspond to multiple reasonable emojis. To better assess a model's ability to align with such real-world emoji usage, we propose a new semantics preserving evaluation framework for emoji recommendation, which measures a model's ability to recommend emojis that maintain the semantic consistency with the user's text. To evaluate how well a model preserves semantics, we assess whether the predicted affective state, demographic profile, and attitudinal stance of the user remain unchanged. If these attributes are preserved, we consider the recommended emojis to have maintained the original semantics. The advanced abilities of Large Language Models (LLMs) in understanding and generating nuanced, contextually relevant output make them well-suited for handling the complexities of semantics preserving emoji recommendation. To this end, we construct a comprehensive benchmark to systematically assess the performance of six proprietary and open-source LLMs using different prompting techniques on our task. Our experiments demonstrate that GPT-4o outperforms other LLMs, achieving a semantics preservation score of 79.23%. Additionally, we conduct case studies to analyze model biases in downstream classification tasks and evaluate the diversity of the recommended emojis.

表情包推荐大模型语义保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。