对比人类与AI生成说服性文本的可检测性,发现微妙的AI persuasion更难被识别。
Can AI-Generated Persuasion Be Detected? Persuaficial Benchmark and AI vs. Human Linguistic Differences
- 构建多语言说服性文本基准Persuaficial,覆盖六种语言。
- 细微的AI生成说服内容使自动检测准确率下降12%以上。
- 揭示人机说服语言差异,助力开发更鲁棒的检测工具。
大型语言模型(LLMs)能生成高度具有说服力的文本,引发其被用于宣传、操纵等有害用途的担忧。本文核心问题为:相较于人类撰写的说服性文本,LLM生成的说服内容是否更难被自动检测?为此,我们分类了可控生成说服性内容的LLM方法,并引入Persuaficial——一个高质量的多语言基准,涵盖英语、德语、波兰语、意大利语、法语和俄语。基于该基准,我们进行了广泛的实证评估,比较了人类与LLM生成的说服性文本。结果表明,虽然明显具有说服力的LLM文本更容易被检测,但细微的LLM生成说服内容会持续降低自动检测性能。此外,我们首次提供了全面的语言学分析,对比人类与LLM生成的说服性文本,为开发更可解释、更鲁棒的检测工具提供洞见。
原文摘要 · Abstract (English)
Large Language Models (LLMs) can generate highly persuasive text, raising concerns about their misuse for propaganda, manipulation, and other harmful purposes. This leads us to our central question: Is LLM-generated persuasion more difficult to automatically detect than human-written persuasion? To address this, we categorize controllable generation approaches for producing persuasive content with LLMs and introduce Persuaficial, a high-quality multilingual benchmark covering six languages: English, German, Polish, Italian, French and Russian. Using this benchmark, we conduct extensive empirical evaluations comparing human-authored and LLM-generated persuasive texts. We find that although overtly persuasive LLM-generated texts can be easier to detect than human-written ones, subtle LLM-generated persuasion consistently degrades automatic detection performance. Beyond detection performance, we provide the first comprehensive linguistic analysis contrasting human and LLM-generated persuasive texts, offering insights that may guide the development of more interpretable and robust detection tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。