提出医疗智能新范式,让AI在不确定中保持谨慎、持续守护临床决策。
Beyond Medical Chatbots: Meddollina and the Rise of Continuous Clinical Intelligence
- 用治理先行机制约束生成,确保推理符合临床规范
- 在1.6万+医疗查询中表现更稳定,减少错误自信与随意推断
- 适合临床辅助系统研发者,关注AI行为而非文本流畅度
生成式医疗AI虽看似流畅且知识丰富,但临床推理远非文本生成。它需在模糊、证据不全和长期上下文下承担责任。当前以生成为中心的系统仍存在过早定论、无根据确信、意图漂移和多步决策不稳定等问题,根源在于将医学视为下一个词预测。本文提出临床情境智能(CCI)作为真实临床应用所需的能力类别,强调持续上下文感知、意图保持、有限推理及证据不足时的合理延迟。我们构建了Meddollina——一种以治理为先的连续临床智能系统,通过在语言生成前约束推理过程,优先保障临床合理性而非生成完整性。该系统作为临床工作流的持续支持层,始终保留医生主导权。在超过16,412个异构医疗查询上评估发现,相比通用模型、医学微调模型及检索增强系统,Meddollina展现出校准的不确定性、对信息不足时的保守推理、长期约束的一致性以及更低的推测性补全。结果表明,可部署的医疗AI不会仅靠规模提升实现,必须转向以临床行为为导向的连续临床智能,其进步应以医生认可的不确定环境下表现衡量,而非文本流畅度。
原文摘要 · Abstract (English)
Generative medical AI now appears fluent and knowledgeable enough to resemble clinical intelligence, encouraging the belief that scaling will make it safe. But clinical reasoning is not text generation. It is a responsibility-bound process under ambiguity, incomplete evidence, and longitudinal context. Even as benchmark scores rise, generation-centric systems still show behaviours incompatible with clinical deployment: premature closure, unjustified certainty, intent drift, and instability across multi-step decisions. We argue these are structural consequences of treating medicine as next-token prediction. We formalise Clinical Contextual Intelligence (CCI) as a distinct capability class required for real-world clinical use, defined by persistent context awareness, intent preservation, bounded inference, and principled deferral when evidence is insufficient. We introduce Meddollina, a governance-first clinical intelligence system designed to constrain inference before language realisation, prioritising clinical appropriateness over generative completeness. Meddollina acts as a continuous intelligence layer supporting clinical workflows while preserving clinician authority. We evaluate Meddollina using a behaviour-first regime across 16,412+ heterogeneous medical queries, benchmarking against general-purpose models, medical-tuned models, and retrieval-augmented systems. Meddollina exhibits a distinct behavioural profile: calibrated uncertainty, conservative reasoning under underspecification, stable longitudinal constraint adherence, and reduced speculative completion relative to generation-centric baselines. These results suggest deployable medical AI will not emerge from scaling alone, motivating a shift toward Continuous Clinical Intelligence, where progress is measured by clinician-aligned behaviour under uncertainty rather than fluency-driven completion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。