arXiv:2604.04743cs.CLcs.AI2026-04被引 2

用几何动力系统解释大模型幻觉,发现幻觉源于任务相关的潜在空间结构。

Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations

  • 将幻觉视为潜在空间中任务依赖的吸引域结构。
  • 事实类任务有清晰吸引域,摘要类任务则常重叠不稳。
  • 通过几何引导可降幻觉率,无需重新训练模型。

大型语言模型会产生幻觉:生成流畅但事实错误的内容。本文提出一种几何动力系统框架,认为幻觉源于任务依赖的潜在空间吸引域结构。通过多个开源模型和基准上的自回归隐藏状态轨迹分析,发现吸引域可分性具有强烈任务依赖性:事实型任务通常呈现更清晰的吸引域分离,而摘要和谬误密集型任务则较不稳定且常发生重叠。我们以任务复杂度与多吸引域定理形式化该行为,刻画了L层Transformer中的吸引域涌现机制,并证明几何感知引导可在不重新训练的情况下降低幻觉概率。

原文摘要 · Abstract (English)

Large language models (LLMs) hallucinate: they produce fluent outputs that are factually incorrect. We present a geometric dynamical systems framework in which hallucinations arise from task-dependent basin structure in latent space. Using autoregressive hidden-state trajectories across multiple open-source models and benchmarks, we find that separability is strongly task-dependent rather than universal: factoid settings can show clearer basin separation, whereas summarization and misconception-heavy settings are typically less stable and often overlap. We formalize this behavior with task-complexity and multi-basin theorems, characterize basin emergence in L-layer transformers, and show that geometry-aware steering can reduce hallucination probability without retraining.

大模型幻觉几何动力系统潜在空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。