arXiv:2604.02637cs.CL2026-04

通过扮演AI角色训练,提升用户对大模型说服力的抵抗力。

Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training

论文配图:Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training
图 1 · 摘自论文原文
  • 让用户扮演LLM角色,亲历预训练、微调和强化学习全过程。
  • 实验显示参与者的抗说服能力显著提升,真理与责任意识增强。
  • 适合想主动应对AI影响的普通用户和教育场景使用。

随着大语言模型(LLMs)日益具有说服力,人们在各种场景中可能受到大规模影响。现有缓解手段(如AI检测器和免责声明)多将人视为被动接收者。为此,我们提出$ extbf{LLMimic}$——一种基于角色扮演的交互式、游戏化AI素养教程,参与者扮演LLM,经历预训练、监督微调(SFT)和人类反馈强化学习(RLHF)三个阶段。我们开展一项$2 \times 3$被试间实验(N=274),参与者要么观看AI历史视频(对照组),要么使用LLMimic(实验组),随后进入三种真实情境的AI说服任务:(a)慈善捐赠劝募,(b)恶意金钱索要,(c)酒店推荐。结果表明,LLMimic显著提升用户AI素养(p < .001),降低各情境下的说服成功率(p < .05),并在酒店推荐情境中显著提高真实性与社会责任感(p < .01)。这表明LLMimic是一种可扩展、以人为本的干预方式,有助于提升公众对说服性AI的认知与应对能力。

原文摘要 · Abstract (English)

As large language models (LLMs) become increasingly persuasive, there is concern that people's opinions and decisions may be influenced across various contexts at scale. Prior mitigation (e.g., AI detectors and disclaimers) largely treats people as passive recipients of AI-generated information. To provide a more proactive intervention against persuasive AI, we introduce $\textbf{LLMimic}$, a role-play-based, interactive, gamified AI literacy tutorial, where participants assume the role of an LLM and progress through three key stages of the training pipeline (pretraining, SFT, and RLHF). We conducted a $2 \times 3$ between-subjects study ($N = 274$) where participants either (1) watched an AI history video (control) or (2) interacted with LLMimic (treatment), and then engaged in one of three realistic AI persuasion scenarios: (a) charity donation persuasion, (b) malicious money solicitation, or (c) hotel recommendation. Our results show that LLMimic significantly improved participants' AI literacy ($p < .001$), reduced persuasion success across scenarios ($p < .05$), and enhanced truthfulness and social responsibility levels ($p<0.01$) in the hotel scenario. These findings suggest that LLMimic offers a scalable, human-centered approach to improving AI literacy and supporting more informed interactions with persuasive AI.

AI素养角色扮演说服力防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。