arXiv:2410.14262cs.CRcs.CL2024-10被引 9

用AI自检自纠,让大模型生成内容更靠谱。

Good Parenting is all you need -- Multi-agentic LLM Hallucination Mitigation

  • 让一个AI写内容,另一个AI检查并修正错误
  • 在4900次测试中,纠错成功率高达85%至100%
  • 适合需要高准确率的AI写作与内容审核场景

本研究探讨大型语言模型(LLM)代理检测并纠正生成内容中幻觉的能力。一个主代理被要求撰写一篇关于虚构丹麦艺术家Flipfloppidy的博客,随后由另一代理审查事实错误。大多数LLM均虚构了该艺术家的存在。在涉及多种主代理与审查代理组合的4900次测试中,Llama3-70b和GPT-4等先进模型在识别幻觉方面表现出近乎完美的准确性,并在收到反馈后成功修正输出,成功率在85%至100%之间。这些结果表明,先进AI模型能显著提升生成内容的准确性和可靠性,为优化AI工作流编排提供了有前景的解决方案。

原文摘要 · Abstract (English)

This study explores the ability of Large Language Model (LLM) agents to detect and correct hallucinations in AI-generated content. A primary agent was tasked with creating a blog about a fictional Danish artist named Flipfloppidy, which was then reviewed by another agent for factual inaccuracies. Most LLMs hallucinated the existence of this artist. Across 4,900 test runs involving various combinations of primary and reviewing agents, advanced AI models such as Llama3-70b and GPT-4 variants demonstrated near-perfect accuracy in identifying hallucinations and successfully revised outputs in 85% to 100% of cases following feedback. These findings underscore the potential of advanced AI models to significantly enhance the accuracy and reliability of generated content, providing a promising approach to improving AI workflow orchestration.

幻觉消除多智能体内容可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。