arXiv:2608.14804cs.AI2026-08

提出临床AI问责的治理框架,区分生成文本与真实病患状态

Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning

论文配图:Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning
图 1 · 摘自论文原文
  • 将临床推理视为部分可观测下的状态估计问题
  • 定义四类信息需求以实现可审计的临床AI
  • 提供六级成熟度模型评估系统治理能力

大型语言模型(LLMs)已成为临床人工智能的主要接口,但其输入输出方式(一次一个上下文窗口)并未显式维护患者当前状态的持久、受控表示。本文认为,纵向临床推理本质上是部分可观测下的状态估计问题,临床AI成败的关键不在于模型读取病历的流畅性,而在于其推理所依赖的患者状态的治理水平。论文区分了生成上下文与受控状态,明确划分了五类常被混淆的对象:真实状态、观测、证据、信念和模拟状态;提出了可审计的分层治理标准;并指出问责的可操作定义分解为四项信息要求:带感知时间版本化的不可变证据账本、与累积证据分离的信念状态、观测过程模型以及声明级因果类型。该分解为分析性框架而非必然定理,核心价值在于概念上的清洁性,使‘可问责临床AI’从口号变为可审计工具。六级成熟度框架区分了系统可治理的范围与可计算的能力,指出当前以LLM为中心的实践虽能力强但成熟度低。本文完全自洽:四个研究问题在引言中明确提出,结论部分逐一回应;未来工作将构建架构核心与完整临床世界模型的研究计划。文中不宣称任何实证结果。

原文摘要 · Abstract (English)

Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in, text out, one context window at a time) maintains no explicit, persistent, governed representation of what is currently true about a patient. This paper argues that longitudinal clinical reasoning is a state-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the governance of the patient state it reasons over. We distinguish generated context from governed state; separate five objects that clinical AI habitually conflates (true state, observations, evidence, belief, and simulated state); define a tiered governance standard against which any clinical AI system can be audited; and show that an operational definition of accountability decomposes into four information requirements: an immutable evidence ledger with awareness-time versioning, a belief state distinct from accumulated evidence, an observation-process model, and claim-level causal typing. We are explicit that this decomposition is analytic rather than a necessity theorem, and that its value is conceptual hygiene: it converts "accountable clinical AI" from a slogan into an audit instrument. A six-level maturity framework separates what a system makes governable from what it can compute, locating current LLM-centric practice at high capability but low maturity. The paper is fully self-contained: the four research questions the framework poses are stated in the introduction, and the conclusion records what the paper establishes toward each; future work develops the buildable core of the architecture and the research program toward full Clinical World Models. No empirical result is claimed here.

临床AI状态治理可问责性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。