量化分析大模型生成中幻觉与创造力的权衡关系。
Shakespearean Sparks: The Dance of Hallucination and Creativity in LLMs' Decoding Layers
- 提出HCL框架,按层评估大模型解码时的幻觉与创造力。
- 发现幻觉与创造力存在稳定权衡,且不同模型有最优平衡层。
- 大模型最优层在早期出现,且此时置信度更高,适合应用。
大语言模型(LLMs)常产生幻觉,这一现象常被认为与创造力相关。以往研究多从理论或定性角度探讨二者关系,本文采用定量方法系统分析该关联。针对创造力的复杂性,我们为LLMs提出一个狭义定义,并构建评估框架HCL,量化解码过程中不同层的幻觉与创造力。实证分析显示,幻觉与创造力的权衡关系在层深、模型类型和模型规模上均保持一致。值得注意的是,不同架构下各模型大小均存在一个最优平衡层;且该层在较大模型中趋向于早期位置,此时模型置信度也显著更高。研究结果为理解大模型创造力与幻觉的相互作用提供了新视角。代码与数据见:https://github.com/ZicongHe2002/HCL-Spark。
原文摘要 · Abstract (English)
Large language models (LLMs) are known to hallucinate, a phenomenon often linked to creativity. While previous research has primarily explored this connection through theoretical or qualitative lenses, our work takes a quantitative approach to systematically examine the relationship between hallucination and creativity in LLMs. Given the complex nature of creativity, we propose a narrow definition tailored to LLMs and introduce an evaluation framework, HCL, which quantifies Hallucination and Creativity across different Layers of LLMs during decoding. Our empirical analysis reveals a tradeoff between hallucination and creativity that is consistent across layer depth, model type, and model size. Notably, across different model architectures, we identify a specific layer at each model size that optimally balances this tradeoff. Additionally, the optimal layer tends to appear in the early layers of larger models, and the confidence of the model is also significantly higher at this layer. These findings provide a quantitative perspective that offers new insights into the interplay between LLM creativity and hallucination. The code and data for our experiments are available at https://github.com/ZicongHe2002/HCL-Spark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。