arXiv:2603.19182cs.AIcs.CL2026-03被引 1

通过分层控制架构提升大模型推理可靠性,减少幻觉。

Box Maze: A Process-Control Architecture for Reliable LLM Reasoning

  • 将推理拆分为记忆锚定、结构化推断和边界控制三层
  • 对抗测试中边界失效率从40%降至1%以下
  • 适合关注大模型安全与可信推理的研究者

大型语言模型(LLMs)虽具备强大生成能力,但在对抗性提示下仍易产生幻觉和不可靠推理。现有安全方法如基于人类反馈的强化学习(RLHF)和输出过滤主要在行为层面操作,缺乏显式的推理过程控制机制。本文提出Box Maze框架,一种概念性的过程控制架构,将LLM推理分解为三个明确层级:记忆锚定、结构化推断和边界强制。我们通过模拟评估,在多个异构LLM系统(DeepSeek-V3、Doubao、Qwen)上进行渐进式边界侵蚀测试,n=50个对抗场景结果表明,显式认知控制层可显著提升边界保持一致性,使边界失效率从基线RLHF的约40%降至对抗条件下的1%以下。尽管当前验证为仿真性质,初步结果表明过程级控制可能是提升大模型推理可靠性的可行方向。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate strong generative capabilities but remain vulnerable to hallucination and unreliable reasoning under adversarial prompting. Existing safety approaches -- such as reinforcement learning from human feedback (RLHF) and output filtering -- primarily operate at the behavioral level and may lack explicit architectural mechanisms for enforcing reasoning process integrity. This paper proposes the Box Maze framework, a conceptual process-control architecture that decomposes LLM reasoning into three explicit layers: memory grounding, structured inference, and boundary enforcement. We introduce preliminary simulation-based evaluation involving progressive boundary erosion scenarios across multiple heterogeneous LLM systems (DeepSeek-V3, Doubao, Qwen). Results from n=50 adversarial scenarios suggest that explicit cognitive control layers may improve consistency in boundary maintenance, with architectural constraints reducing boundary failure rates from approximately 40% (baseline RLHF) to below 1% under adversarial conditions. While current validation is simulation-based, these preliminary results indicate that process-level control may offer a promising direction for improving reliability in large language model reasoning.

大模型推理过程控制可靠性幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。