构建58种社交影响技术分类体系,评测大模型识别文本隐性操控能力。
Unraveling SITT: Social Influence Technique Taxonomy and Detection with LLMs
- 提出9类58种实证驱动的社交影响技术分类框架
- 大模型在识别上下文敏感技巧时表现有限,最优F1仅0.45
- 适合研究虚假信息、舆论操控与AI伦理的学者使用
本文提出社交影响技术分类体系(SITT),包含58种基于实证的技术,分为九类,旨在检测文本中微妙的社交影响。基于跨学科基础,构建了746条对话组成的SITT数据集,由11位波兰语专家标注并翻译成英文,用于评估大模型识别这些技术的能力。采用分层多标签分类方法,对GPT-4o、Claude 3.5、Llama-3.1、Mixtral和PLLuM五款模型进行基准测试。结果显示,尽管Claude 3.5在类别识别上表现较好(F1=0.45),但整体性能受限,尤其在处理上下文敏感技巧时表现不佳。研究揭示当前大模型对细微语言线索的感知能力不足,强调领域特定微调的重要性。该工作为理解大模型如何检测、分类甚至复制自然对话中的社交影响策略提供了新资源与评估范例。
原文摘要 · Abstract (English)
In this work we present the Social Influence Technique Taxonomy (SITT), a comprehensive framework of 58 empirically grounded techniques organized into nine categories, designed to detect subtle forms of social influence in textual content. We also investigate the LLMs ability to identify various forms of social influence. Building on interdisciplinary foundations, we construct the SITT dataset -- a 746-dialogue corpus annotated by 11 experts in Polish and translated into English -- to evaluate the ability of LLMs to identify these techniques. Using a hierarchical multi-label classification setup, we benchmark five LLMs, including GPT-4o, Claude 3.5, Llama-3.1, Mixtral, and PLLuM. Our results show that while some models, notably Claude 3.5, achieved moderate success (F1 score = 0.45 for categories), overall performance of models remains limited, particularly for context-sensitive techniques. The findings demonstrate key limitations in current LLMs' sensitivity to nuanced linguistic cues and underscore the importance of domain-specific fine-tuning. This work contributes a novel resource and evaluation example for understanding how LLMs detect, classify, and potentially replicate strategies of social influence in natural dialogues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。