arXiv:2512.21635cs.CL2025-12KDD被引 2

区分智能幻觉与缺陷幻觉,揭示大模型幻想背后的创造力潜力。

Heaven-Sent or Hell-Bent? Benchmarking the Intelligence and Defectiveness of LLM Hallucinations

  • 按创意性与真实性将幻觉分类,构建多维评估体系。
  • 实验证明幻觉与创造力存在非线性关系,可协同优化。
  • 适合研究大模型创新力、可信度的学者与工程师参考。

大语言模型(LLM)中的幻觉通常被视为需最小化的错误,但新观点认为部分幻觉可能蕴含创造性或认识论价值,这一维度在现有研究中尚未被量化。当前幻觉检测方法多聚焦事实一致性,难以应对多样化的科学任务,且难平衡创意与准确性。为此,我们提出HIC-Bench——一个将幻觉分为智能幻觉(IH)与缺陷幻觉(DH)的新评估框架,系统研究其在模型创造力中的相互作用。该框架具备三大特性:(1) 结构化评估:融合托兰斯创造力测试(TTCT)指标(原创性、可行性、价值)与幻觉特异性维度(科学合理性、事实偏离度);(2) 跨领域适用:覆盖十大学科领域,涵盖开放式创新任务;(3) 动态提示优化:通过动态幻觉提示(DHP)引导模型生成既具创意又可靠的内容。评估采用多轮LLM裁判打分并取均值以降低偏见,人工标注验证分类结果。实验表明,IH与DH间存在非线性关系,证明创造力与正确性可共同优化。这些发现将IH定位为创造力催化剂,揭示了大模型幻觉驱动科学创新的能力。同时,HIC-Bench为深入探索大模型幻觉的创造性智能提供了重要平台。

原文摘要 · Abstract (English)

Hallucinations in large language models (LLMs) are commonly regarded as errors to be minimized. However, recent perspectives suggest that some hallucinations may encode creative or epistemically valuable content, a dimension that remains underquantified in current literature. Existing hallucination detection methods primarily focus on factual consistency, struggling to handle heterogeneous scientific tasks and balance creativity with accuracy. To address these challenges, we propose HIC-Bench, a novel evaluation framework that categorizes hallucinations into Intelligent Hallucinations (IH) and Defective Hallucinations (DH), enabling systematic investigation of their interplay in LLM creativity. HIC-Bench features three core characteristics: (1) Structured IH/DH Assessment. using a multi-dimensional metric matrix integrating Torrance Tests of Creative Thinking (TTCT) metrics (Originality, Feasibility, Value) with hallucination-specific dimensions (scientific plausibility, factual deviation); (2) Cross-Domain Applicability. spanning ten scientific domains with open-ended innovation tasks; and (3) Dynamic Prompt Optimization. leveraging the Dynamic Hallucination Prompt (DHP) to guide models toward creative and reliable outputs. The evaluation process employs multiple LLM judges, averaging scores to mitigate bias, with human annotators verifying IH/DH classifications. Experimental results reveal a nonlinear relationship between IH and DH, demonstrating that creativity and correctness can be jointly optimized. These insights position IH as a catalyst for creativity and reveal the ability of LLM hallucinations to drive scientific innovation.Additionally, the HIC-Bench offers a valuable platform for advancing research into the creative intelligence of LLM hallucinations.

大模型幻觉创造力评估智能评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。