arXiv:2412.02790cs.CLcs.AI2024-12被引 2

用进化算法生成更准的问答数据,减少大模型胡说八道。

An Evolutionary Large Language Model for Hallucination Mitigation

  • 模仿生物进化,用遗传算法优化问答对生成
  • 在深度、相关性、覆盖率上优于人工数据集
  • 适合需要高准确率的医疗法律等领域

大型语言模型(如ChatGPT、Gemini)推动了AI应用的新时代,能生成文本、图像和视频。但这些模型常出现幻觉问题——自信地输出错误或虚构信息。在医疗、法律等专业领域,信息准确性至关重要。本文提出EvoLLMs,一种受进化计算启发的框架,自动生成高质量问答数据集,有效降低幻觉。该框架采用遗传算法,模拟选择、变异、突变等过程,引导大模型生成准确且语境相关的问答对。对比分析显示,EvoLLMs在深度、相关性和覆盖度等关键指标上持续优于人工生成数据集,同时在抑制幻觉方面几乎达到人类水平。结果表明,EvoLLMs是高效可靠的问答数据生成方案,显著减少人工标注的时间与资源消耗。

原文摘要 · Abstract (English)

The emergence of LLMs, like ChatGPT and Gemini, has marked the modern era of artificial intelligence applications characterized by high-impact applications generating text, images, and videos. However, these models usually ensue with one critical challenge called hallucination: confident presentation of inaccurate or fabricated information. This problem attracts serious concern when these models are applied to specialized domains, including healthcare and law, where the accuracy and preciseness of information are absolute conditions. In this paper, we propose EvoLLMs, an innovative framework inspired by Evolutionary Computation, which automates the generation of high-quality Question-answering (QA) datasets while minimizing hallucinations. EvoLLMs employs genetic algorithms, mimicking evolutionary processes like selection, variation, and mutation, to guide LLMs in generating accurate, contextually relevant question-answer pairs. Comparative analysis shows that EvoLLMs consistently outperforms human-generated datasets in key metrics such as Depth, Relevance, and Coverage, while nearly matching human performance in mitigating hallucinations. These results highlight EvoLLMs as a robust and efficient solution for QA dataset generation, significantly reducing the time and resources required for manual curation.

大模型幻觉抑制生成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。