arXiv:2604.14881cs.AIcs.CY2026-04被引 3

让AI和人一起理性决策,关键在识别并稳定思维漂移。

The Missing Knowledge Layer in AI: A Framework for Stable Human-AI Reasoning

  • 构建双层机制:人类侧设不确定提示、冲突暴露,模型侧用认知控制环检测不稳
  • 使推理过程可追溯,提前暴露不确定性,避免高风险误判
  • 适合关注AI治理、可信决策的从业者,尤其合规场景

大语言模型在医疗、法律、金融、工程和政府等领域的决策中日益普及,但其核心缺陷在于:即使内部推理已偏离,仍能生成流畅输出。自信的回答可能隐藏不确定性、猜测或不一致,微小表述变化即可导致结论迥异。这使模型成为有用助手,却难以作为高风险场景中的可靠伙伴。人类也存在类似弱点,常将流畅度误认为可靠性。当模型回答顺滑时,用户容易盲信,双方可能共同漂移。本文是五篇系列研究的第一篇,提出双层稳定人机推理框架:第二至第四篇引入人类侧机制,如不确定提示、冲突揭示与可审计推理路径;第五篇则开发模型侧的表征控制环(Epistemic Control Loop, ECL),实时检测推理不稳定性并调节生成。二者协同构成缺失的治理底层,提升使用场景下的信号噪声比。通过提前暴露不确定性与漂移,实现更精准的能力管控,契合欧盟《人工智能法案》和ISO/IEC 42001等新兴合规要求,使推理过程在真实使用条件下具备可追溯性。核心主张是:流畅≠可靠。若缺乏同时稳定人与模型推理的结构,AI在关键领域无法被信任或有效治理。

原文摘要 · Abstract (English)

Large language models are increasingly integrated into decision-making in areas such as healthcare, law, finance, engineering, and government. Yet they share a critical limitation: they produce fluent outputs even when their internal reasoning has drifted. A confident answer can conceal uncertainty, speculation, or inconsistency, and small changes in phrasing can lead to different conclusions. This makes LLMs useful assistants but unreliable partners in high-stakes contexts. Humans exhibit a similar weakness, often mistaking fluency for reliability. When a model responds smoothly, users tend to trust it, even when both model and user are drifting together. This paper is the first in a five-paper research series on stabilising human-AI reasoning. The series proposes a two-layer approach: Parts II-IV introduce human-side mechanisms such as uncertainty cues, conflict surfacing, and auditable reasoning traces, while Part V develops a model-side Epistemic Control Loop (ECL) that detects instability and modulates generation accordingly. Together, these layers form a missing operational substrate for governance by increasing signal-to-noise at the point of use. Stabilising interaction makes uncertainty and drift visible before enforcement is applied, enabling more precise capability governance. This aligns with emerging compliance expectations, including the EU AI Act and ISO/IEC 42001, by making reasoning processes traceable under real conditions of use. The central claim is that fluency is not reliability. Without structures that stabilise both human and model reasoning, AI cannot be trusted or governed where it matters most.

AI治理人机协作推理稳定可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。