三种去幻觉方法对大模型创造力影响各异,科学应用需权衡准确与创新。
Does Less Hallucination Mean Less Creativity? An Empirical Investigation in LLMs
- 对比验证、层间对比解码和检索增强生成三类去幻觉方法
- 链式验证提升发散思维,层间对比抑制创造力,检索增强影响小
- 适合科研场景中需兼顾事实准确与创意探索的用户参考
大型语言模型在自然语言理解与推理方面表现卓越,但存在生成虚假内容的幻觉问题。尽管已有多种去幻觉方法,其对创造性生成的影响仍不明朗,尤其在需要兼具事实准确性与创意假设生成的AI辅助科学发现中尤为关键。本文研究了三种去幻觉技术——链式验证(CoVe)、通过对比层解码(DoLa)和检索增强生成(RAG)——对大模型创造力的影响。我们在多个模型家族(LLaMA、Qwen、Mistral)及不同规模(1B–70B参数)下,基于两个创造力评测基准(NeoCoder 和 CS4)进行评估,发现这些方法对发散性创造力有相反影响:CoVe 提升发散思维,DoLa 抑制创造力,而 RAG 影响不显著。结果为科学应用中选择合适的去幻觉方法提供了依据,强调在事实准确与创意探索之间取得平衡的重要性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) exhibit remarkable capabilities in natural language understanding and reasoning, but suffer from hallucination: the generation of factually incorrect content. While numerous methods have been developed to reduce hallucinations, their impact on creative generations remains unexplored. This gap is particularly critical for AI-assisted scientific discovery, which requires both factual accuracy and creative hypothesis generation. We investigate how three hallucination-reduction techniques: Chain of Verification (CoVe), Decoding by Contrasting Layers (DoLa), and Retrieval-Augmented Generation (RAG), affect creativity in LLMs. Evaluating multiple model families (LLaMA, Qwen, Mistral) at varying scales (1B - 70B parameters) on two creativity benchmarks (NeoCoder and CS4), we find that these methods have opposing effects on divergent creativity. CoVe enhances divergent thinking, DoLa suppresses it, and RAG shows minimal impact. Our findings provide guidance for selecting appropriate hallucination-reduction methods in scientific applications, where the balance between factual accuracy and creative exploration is crucial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。