arXiv:2608.27484cs.AIcs.IR2026-08

CareGraph将多元健康数据转为可审计的智能洞察,提升医疗决策透明度。

CareGraph: An Auditable Hybrid AI Framework for Evidence-Grounded Personalized Longitudinal Health Intelligence

论文配图:CareGraph: An Auditable Hybrid AI Framework for Evidence-Grounded Personalized Longitudinal Health Intelligence
图 1 · 摘自论文原文
  • 构建混合框架,统一临床、自报与可穿戴数据,生成趋势与上下文缺失提示
  • 在400人测试集上实现0.837宏F1,上下文缺失检测性能超旧方法2.6倍
  • 适合医疗AI系统开发者,提供安全可控、可追溯的个性化健康分析基础

人工智能正重塑个性化医疗,但临床、自报及可穿戴数据碎片化,难以解释与追踪。本文提出CareGraph,一种可审计的混合AI框架,将异构记录转化为优先级趋势、缺失上下文标识、边界化下一步建议、讨论问题及溯源链接的解释。CareGraph仅组织证据,不进行诊断、预测、治疗选择或自主决策。其流程涵盖确定性分析、上下文检测、图结构构建、受限语言模型合成、证据验证、安全控制与发布门控。使用每组400名患者的合成队列进行开发、验证与留出测试。在留出数据上,冻结的普通最小二乘趋势规则配合充分性门控达到0.827准确率、0.837宏F1(95%置信区间0.819至0.854),不足数据F1达0.974。缺失上下文检测实现0.815严格微F1,较旧检测器(0.318)显著提升。在人工编写的留出基准上,安全规则集1.2版达成1.000精确率、0.950召回率和0.974 F1。一次涉及80名患者的审计中,成功生成79份合成结果并呈现78份,仅1例因无效证据键被阻断,1例失败关闭。与单体GPT-5.6在56名匹配患者上的对比显示,CareGraph耗时40.15秒优于49.62秒,输出长度661词短于1,163词,且与纵向目标的探索性词汇对齐更优;基线模型用更少令牌,引用更多原始证据。图审计验证了溯源与确定性检索;图增量对生成的影响需配对评估。CareGraph为智能个性化健康系统提供了安全约束的基础。

原文摘要 · Abstract (English)

Artificial intelligence is transforming personalized healthcare, yet fragmented clinical, self reported, and wearable evidence remains difficult to interpret and trace. We present CareGraph, an auditable hybrid AI framework that converts heterogeneous records into prioritized trends, missing context indicators, bounded next steps, discussion questions, and provenance linked explanations. CareGraph organizes evidence without diagnosing, predicting outcomes, selecting treatment, or making autonomous clinical decisions. Its pipeline covers deterministic analysis, context detection, graph construction, constrained language model synthesis, evidence validation, safety controls, and release gating. Tests used synthetic cohorts of 400 patients each for development, validation, and holdout. On holdout data, a frozen ordinary least squares trend rule with a sufficiency gate achieved 0.827 accuracy, 0.837 macro F1 with a 95 percent confidence interval of 0.819 to 0.854, and 0.974 insufficient data F1. Missing context detection achieved 0.815 strict micro F1 versus 0.318 for the legacy detector. On an authored holdout benchmark, safety ruleset version 1.2 achieved 1.000 precision, 0.950 recall, and 0.974 F1. An audit requiring graph retrieval across 80 patients yielded 79 syntheses and 78 presentations without fallback; one output was blocked and one failed closed because of an invalid evidence key. Against monolithic GPT 5.6 on 56 matched patients, CareGraph was faster at 40.15 versus 49.62 seconds, shorter at 661 versus 1,163 words, and showed better exploratory lexical alignment with longitudinal targets; the baseline used fewer tokens and cited more raw evidence. Graph auditing verified provenance and deterministic retrieval; incremental graph effects on generation require paired evaluation. CareGraph offers a safety bounded foundation for intelligent personalized health systems.

医疗AI可审计性健康智能多源数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。