arXiv:2508.15798cs.CLcs.AI2025-08被引 1

研究大模型如何用话术说服人并放大偏见,揭示其潜在滥用风险。

Persuasiveness and Bias in LLM: Investigating the Impact of Persuasiveness and Reinforcement of Bias in Language Models

  • 用角色扮演框架模拟说服者与质疑者,量化信念改变程度
  • 发现强说服力模型会无意中强化种族、性别、宗教偏见
  • 适合关注AI伦理、内容安全与可信部署的研究者和开发者

大型语言模型(LLMs)能生成逼真、类人的文本,广泛应用于内容创作、决策支持与用户交互。但同样的系统可能大规模传播信息或虚假信息,并反映数据、架构或训练选择带来的社会偏见。本研究探讨说服力与偏见之间的相互作用,重点分析不完美或有偏差的输出如何影响说服效果。具体而言,测试基于角色的模型是否能在使用事实性主张的同时,无意中推广虚假信息或偏见叙事。我们引入说服者-质疑者框架:说服者模型采用角色模拟真实态度,质疑者模型作为人类代理,比较其在接触说服者论点前后信念分布的变化。说服力通过信念分布的Jensen-Shannon散度进行量化。随后考察被说服实体是否会进一步强化并放大涉及种族、性别和宗教的偏见。对强说服者模型使用谄媚式对抗提示进行深入探测,并由其他模型评估其偏见水平。结果表明,尽管大模型具备塑造叙事、调整语调、匹配受众价值观的能力,可应用于心理学、营销和法律援助等领域,但其能力也可能被滥用,用于自动化制造虚假信息或设计利用认知偏见的攻击性话语,加剧刻板印象与不平等。核心风险在于恶意使用而非偶发错误。通过测量说服力与偏见强化,我们主张建立防御机制与政策,惩罚欺骗性使用,推动对齐、价值敏感设计及可信部署。

原文摘要 · Abstract (English)

Warning: This research studies AI persuasion and bias amplification that could be misused; all experiments are for safety evaluation. Large Language Models (LLMs) now generate convincing, human-like text and are widely used in content creation, decision support, and user interactions. Yet the same systems can spread information or misinformation at scale and reflect social biases that arise from data, architecture, or training choices. This work examines how persuasion and bias interact in LLMs, focusing on how imperfect or skewed outputs affect persuasive impact. Specifically, we test whether persona-based models can persuade with fact-based claims while also, unintentionally, promoting misinformation or biased narratives. We introduce a convincer-skeptic framework: LLMs adopt personas to simulate realistic attitudes. Skeptic models serve as human proxies; we compare their beliefs before and after exposure to arguments from convincer models. Persuasion is quantified with Jensen-Shannon divergence over belief distributions. We then ask how much persuaded entities go on to reinforce and amplify biased beliefs across race, gender, and religion. Strong persuaders are further probed for bias using sycophantic adversarial prompts and judged with additional models. Our findings show both promise and risk. LLMs can shape narratives, adapt tone, and mirror audience values across domains such as psychology, marketing, and legal assistance. But the same capacity can be weaponized to automate misinformation or craft messages that exploit cognitive biases, reinforcing stereotypes and widening inequities. The core danger lies in misuse more than in occasional model mistakes. By measuring persuasive power and bias reinforcement, we argue for guardrails and policies that penalize deceptive use and support alignment, value-sensitive design, and trustworthy deployment.

大模型伦理偏见放大说服力评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。