arXiv:2410.19485cs.CL2024-10被引 3

让AI互相辩论,发现对抗虚假信息能提升真实回答准确率

A Debate-Driven Experiment on LLM Hallucinations and Accuracy

  • 用多个GPT-4o-Mini模型模拟辩论,一说假话一说真话
  • 在TruthfulQA数据集上,真话方因反驳假话而准确率提升
  • 适合研究AI可信度、模型协作机制的学者参考

大型语言模型(LLMs)虽能生成连贯且符合语境的文本,但仍易产生不基于输入或外部知识的幻觉。以往缓解幻觉的方法包括在高质量数据集上微调模型、引入事实核查机制和对抗训练。这些方法多聚焦于单个模型输出,未探索模型间互动的影响。本研究通过新实验框架,让多个GPT-4o-Mini模型基于TruthfulQA数据集的问题进行类辩论互动:一个模型被指令生成看似合理但错误的答案,其余模型则需提供真实回答。实验旨在评估误导性信息是否促使真实回答方更充分论证自身推理,从而提升在TruthfulQA基准上的表现。结果表明,模型间的交互可为提高输出准确性与鲁棒性提供重要启示,补充现有缓解策略。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved a degree of success in generating coherent and contextually relevant text, yet they remain prone to a significant challenge known as hallucination: producing information that is not substantiated by the input or external knowledge. Previous efforts to mitigate hallucinations have focused on techniques such as fine-tuning models on high-quality datasets, incorporating fact-checking mechanisms, and developing adversarial training methods. While these approaches have shown some promise, they often address the issue at the level of individual model outputs, leaving unexplored the effects of inter-model interactions on hallucination. This study investigates the phenomenon of hallucination in LLMs through a novel experimental framework where multiple instances of GPT-4o-Mini models engage in a debate-like interaction prompted with questions from the TruthfulQA dataset. One model is deliberately instructed to generate plausible but false answers while the other models are asked to respond truthfully. The experiment is designed to assess whether the introduction of misinformation by one model can challenge the truthful majority to better justify their reasoning, improving performance on the TruthfulQA benchmark. The findings suggest that inter-model interactions can offer valuable insights into improving the accuracy and robustness of LLM outputs, complementing existing mitigation strategies.

幻觉缓解模型辩论可信生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。