小模型也能靠高信心说服大模型说谎,研究如何量化这种误导风险。
When Persuasion Overrides Truth in Multi-Agent LLM Debates: Introducing a Confidence-Weighted Persuasion Override Rate (CW-POR)
- 用多智能体辩论框架测试模型被说服程度,结合置信度加权。
- 30-300词的冗长辩词让3B-14B模型常误信错误答案,且信心极高。
- 适用于评估模型抗误导能力,尤其关注高置信错误的风险。
在真实场景中,单个大语言模型可能面临真假并存的冲突陈述,需判断真伪。本文在单轮多智能体辩论框架下研究此风险:一个智能体提供来自TruthfulQA的真实答案,另一个则强烈捍卫虚假结论,同一模型架构作为裁判。我们提出置信度加权说服覆盖率(CW-POR),不仅衡量裁判被误导频率,还反映其对错误选择的置信强度。在五个开源模型(3B-14B参数)上系统测试不同辩词长度(30-300词)的结果显示,即使小型模型也能生成极具说服力的论据,使模型以高置信度采纳错误答案。该发现强调了模型校准与对抗测试的重要性,防止模型自信地传播错误信息。
原文摘要 · Abstract (English)
In many real-world scenarios, a single Large Language Model (LLM) may encounter contradictory claims-some accurate, others forcefully incorrect-and must judge which is true. We investigate this risk in a single-turn, multi-agent debate framework: one LLM-based agent provides a factual answer from TruthfulQA, another vigorously defends a falsehood, and the same LLM architecture serves as judge. We introduce the Confidence-Weighted Persuasion Override Rate (CW-POR), which captures not only how often the judge is deceived but also how strongly it believes the incorrect choice. Our experiments on five open-source LLMs (3B-14B parameters), where we systematically vary agent verbosity (30-300 words), reveal that even smaller models can craft persuasive arguments that override truthful answers-often with high confidence. These findings underscore the importance of robust calibration and adversarial testing to prevent LLMs from confidently endorsing misinformation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。