arXiv:2409.16710cs.CEcs.CL2024-09被引 1

GPT-4生成文本能影响专家判断,揭示大模型决策操控风险

Beyond Turing Test: Can GPT-4 Sway Experts' Decisions?

  • 用真实读者反应评估GPT-4文本说服力,突破传统真假辨识标准
  • 专家与普通人均受生成内容影响,逻辑与说服力是关键影响因素
  • 提出以读者决策反应为评价新范式,适合关注可信度与伦理的研究者

在后图灵时代,评估大语言模型(LLMs)不再仅看生成文本是否像真人写作,而是考察读者的真实反应。本文研究GPT-4生成文本对业余和专业读者决策的影响。结果表明,该模型可生成具有说服力的分析,显著影响两类人群的判断。我们从语法、说服力、逻辑连贯性和实用性四个维度评估生成内容,发现现实受众反应与当前多维评价指标高度相关。本研究揭示了生成文本操纵人类决策的潜力与风险,并提出以读者反应和决策行为作为新评估方向。我们已公开数据集以支持后续研究。

原文摘要 · Abstract (English)

In the post-Turing era, evaluating large language models (LLMs) involves assessing generated text based on readers' reactions rather than merely its indistinguishability from human-produced content. This paper explores how LLM-generated text impacts readers' decisions, focusing on both amateur and expert audiences. Our findings indicate that GPT-4 can generate persuasive analyses affecting the decisions of both amateurs and professionals. Furthermore, we evaluate the generated text from the aspects of grammar, convincingness, logical coherence, and usefulness. The results highlight a high correlation between real-world evaluation through audience reactions and the current multi-dimensional evaluators commonly used for generative models. Overall, this paper shows the potential and risk of using generated text to sway human decisions and also points out a new direction for evaluating generated text, i.e., leveraging the reactions and decisions of readers. We release our dataset to assist future research.

大模型评估说服力决策影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。