arXiv:2604.20055cs.AIcs.HC2026-04

用AI把模糊的医院质量改进过程变规范,提升效率与可审计性。

From Fuzzy to Formal: Scaling Hospital Quality Improvement with AI

论文配图:From Fuzzy to Formal: Scaling Hospital Quality Improvement with AI
图 1 · 摘自论文原文
  • 通过人机协同优化自然语言规范与AI流程,实现探索性分析的标准化。
  • 在真实医院数据上达到超70%专家标注一致性,效率远超传统方法。
  • 适合医疗质量改进团队、临床研究者及希望提升决策透明度的机构。

医院质量改进(QI)对优化医疗交付至关重要,其核心是识别可干预的关键因素,即QI因子发现。传统方法依赖专家主导的半结构化定性工具,如鱼骨图、病历审查和精益医疗法,存在耗时长、资源密集、可重复性差和难追溯等问题。当前的AI对齐方法假设任务定义清晰,但QI因子发现本质上是探索性、模糊且迭代的认知过程,依赖复杂隐含的专家判断。为此,我们提出将该任务视为学习大模型提示词与整体自然语言规范的过程,将其映射到经典的AI/ML开发流程(问题形式化、模型学习、模型验证),其中规范作为可调超参数。领域专家与AI代理共同迭代优化规范与AI流程,直至AI提取结果与专家标注一致并契合临床目标。我们在一家城市安全网医院应用该‘人机规范-解决方案共优化’框架,识别长期住院与30天非计划再入院的关键驱动因素。最终的AI-QI流程达到≥70%与专家标注的一致性。相比以往手动精益分析,该流程更高效,复现了已有发现,揭示了新可干预因素,并生成可审计的推理路径。

原文摘要 · Abstract (English)

Hospital Quality Improvement (QI) plays a critical role in optimizing healthcare delivery by translating high-level hospital goals into actionable solutions. A critical step of QI is to identify the key modifiable contributing factors, a process we call QI factor discovery, typically through expert-driven semi-structured qualitative tools like fishbone diagrams, chart reviews, and Lean Healthcare methods. AI has the potential to transform and accelerate QI factor discovery, which is traditionally time- and resource-intensive and limited in reproducibility and auditability. Nevertheless, current AI alignment methods assume the task is well-defined, whereas QI factor discovery is an exploratory, fuzzy, and iterative sense-making process that relies on complex implicit expert judgments. To design an AI pipeline that formalizes the QI process while preserving its exploratory components, we propose viewing the task as learning not only LLM prompts but also the overarching natural-language specifications. In particular, we map QI factor discovery to steps of the classical AI/ML development process (problem formalization, model learning, and model validation) where the specifications are tunable hyperparameters. Domain experts and AI agents iteratively refine both the overarching specifications and AI pipeline until AI extractions are concordant with expert annotations and aligned with clinical objectives. We applied this "Human-AI Spec-Solution Co-optimization" framework at an urban safety-net hospital to identify factors driving prolonged length of stay and unplanned 30-day readmissions. The resulting AI-for-QI pipelines achieved $\ge 70\%$ concordance with expert annotations. Compared to prior manual Lean analyses, the AI pipeline was substantially more efficient, recovered previous findings, surfaced new modifiable factors, and produced auditable reasoning traces.

医疗AI质量改进人机协同可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。