arXiv:2410.13517cs.CLcs.AI2024-10ACL被引 11

让大模型自己和自己辩论,测试其偏见是否容易被攻破

Bias in the Mirror: Are LLMs opinions robust to their own adversarial attacks ?

  • 用同一模型的两个实例互相辩论对立观点
  • 发现大模型偏见在对抗中仍较顽固,部分会受误导
  • 适合研究模型安全与偏见机制的学者参考

大型语言模型(LLMs)从训练数据和对齐过程继承偏见,以微妙方式影响其回应。尽管已有大量研究关注这些偏见,但对其在交互过程中鲁棒性的探讨仍较少。本文提出一种新方法:让两个同源的LLM实例进行自我辩论,分别持对立立场以说服一个中立版本的模型。通过该机制,评估偏见的稳定性以及模型是否易被误导至传播错误信息或转向有害观点。实验涵盖多种不同规模、来源及语言的LLM,揭示了偏见在跨语言与文化背景下的持续性与可塑性差异。

原文摘要 · Abstract (English)

Large language models (LLMs) inherit biases from their training data and alignment processes, influencing their responses in subtle ways. While many studies have examined these biases, little work has explored their robustness during interactions. In this paper, we introduce a novel approach where two instances of an LLM engage in self-debate, arguing opposing viewpoints to persuade a neutral version of the model. Through this, we evaluate how firmly biases hold and whether models are susceptible to reinforcing misinformation or shifting to harmful viewpoints. Our experiments span multiple LLMs of varying sizes, origins, and languages, providing deeper insights into bias persistence and flexibility across linguistic and cultural contexts.

大模型偏见自我辩论模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。