arXiv:2410.14596cs.CLcs.AI2024-10NAACL被引 21

让大模型学会识别并平衡有害与有益说服,提升判断力与协作稳定性。

Teaching Models to Balance Resisting and Accepting Persuasion

  • 通过多智能体递归对话生成数据,用偏好优化训练模型辨别说服正负
  • 70B模型在双场景下既抗误导又善吸收建议,整体表现最优
  • 显著改善团队辩论中强弱模型的协作稳定性,减少回答顺序依赖

大型语言模型易受说服影响,面对对抗性对话可能带来风险。本文首次提出兼顾抵御负面说服与接受正面说服的防御机制:仅优化单一方向会导致另一侧表现不佳。为此,提出说服训练(PBT),利用小型7-8B模型生成多智能体递归对话数据,通过偏好优化训练大模型(70B)在适当时机接受或拒绝说服。PBT使小模型生成的数据可用于训练大模型,并在包含正负说服的综合数据集上实现最佳整体性能,显著提升对虚假信息的抵抗能力与被挑战时的韧性。更重要的是,在趣味问答和常识推理两个领域,使用PBT的多智能体团队表现更稳定,强模型能持续带动弱模型,打破传统中因回答顺序导致性能波动的问题。

原文摘要 · Abstract (English)

Large language models (LLMs) are susceptible to persuasion, which can pose risks when models are faced with an adversarial interlocutor. We take a first step towards defending models against persuasion while also arguing that defense against adversarial (i.e. negative) persuasion is only half of the equation: models should also be able to accept beneficial (i.e. positive) persuasion to improve their answers. We show that optimizing models for only one side results in poor performance on the other. In order to balance positive and negative persuasion, we introduce Persuasion-Training (or PBT), which leverages multi-agent recursive dialogue trees to create data and trains models via preference optimization to accept persuasion when appropriate. PBT allows us to use data generated from dialogues between smaller 7-8B models for training much larger 70B models. Moreover, PBT consistently improves resistance to misinformation and resilience to being challenged while also resulting in the best overall performance on holistic data containing both positive and negative persuasion. Crucially, we show that PBT models are better teammates in multi-agent debates across two domains (trivia and commonsense QA). We find that without PBT, pairs of stronger and weaker models have unstable performance, with the order in which the models present their answers determining whether the team obtains the stronger or weaker model's performance. PBT leads to better and more stable results and less order dependence, with the stronger model consistently pulling the weaker one up.

大模型说服对抗多智能体协作增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。