arXiv:2410.02653cs.CLcs.CV2024-10被引 21

首个大规模评测框架,量化大模型的说服力并揭示训练策略比规模更重要。

Measuring and Improving Persuasiveness of Large Language Models

  • 构建自动评测框架PersuasionBench与PersuasionArena,覆盖多种说服任务。
  • 模型规模越大越有说服力,但小模型经针对性训练可超越大模型。
  • 强调训练方法比算力规模更关键,为政策制定提供新视角。

大语言模型在生成人类消费内容(如营销文案)和直接与人交互(如聊天机器人)中的应用日益广泛。能够生成可验证说服力内容的系统既带来机遇,也引发社会挑战。为引导其社会影响,亟需建立衡量与基准测试其说服力的方法。为此,本文提出PersuasionBench与PersuasionArena,这是首个大规模自动化评测集,包含多类任务以评估生成模型的说服能力。研究发现,模型说服力与规模正相关,但小模型通过合成与真实数据的定向训练,可显著提升说服力,甚至超越更大模型。这一结果挑战了“规模决定一切”的假设。此外,当前以浮点运算量为标准的监管(如欧盟AI法案、加州SB-1047)无法全面反映模型的社会影响。我们呼吁社区共同参与该平台,推动对人工智能说服力及其社会影响的理解。

原文摘要 · Abstract (English)

LLMs are increasingly being used in workflows involving generating content to be consumed by humans (e.g., marketing) and also in directly interacting with humans (e.g., through chatbots). The development of such systems that are capable of generating verifiably persuasive messages presents both opportunities and challenges for society. On the one hand, such systems could positively impact domains like advertising and social good, such as addressing drug addiction, and on the other, they could be misused for spreading misinformation and shaping political opinions. To channel LLMs' impact on society, we need to develop systems to measure and benchmark their persuasiveness. With this motivation, we introduce PersuasionBench and PersuasionArena, the first large-scale benchmark and arena containing a battery of tasks to measure the persuasion ability of generative models automatically. We investigate to what extent LLMs know and leverage linguistic patterns that can help them generate more persuasive language. Our findings indicate that the persuasiveness of LLMs correlates positively with model size, but smaller models can also be made to have a higher persuasiveness than much larger models. Notably, targeted training using synthetic and natural datasets significantly enhances smaller models' persuasive capabilities, challenging scale-dependent assumptions. Our findings carry key implications for both model developers and policymakers. For instance, while the EU AI Act and California's SB-1047 aim to regulate AI models based on the number of floating point operations, we demonstrate that simple metrics like this alone fail to capture the full scope of AI's societal impact. We invite the community to explore and contribute to PersuasionArena and PersuasionBench, available at https://bit.ly/measure-persuasion, to advance our understanding of AI-driven persuasion and its societal implications.

说服力评测大模型社会影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。