用可计算证据链检测糖尿病模型输出,揪出不靠谱的医疗建议。
T2D-Bench: Evidence-Gated Evaluation of LLM Outputs for Type 2 Diabetes Using a Multi-Layer Clinical-Lifestyle Knowledge Graph
- 构建多层临床生活方式知识图谱,连接医学数据与血糖机制。
- 35%的GPT-4o-mini输出和33%的GPT-4o输出未通过证据验证。
- 支持修正建议,让模型输出可追溯、可纠错,适合医疗AI评估者。
大型语言模型(LLMs)虽能生成看似专业的2型糖尿病建议,却常违反指南要求或无法合理解释生活方式对血糖的影响。我们提出T2D-Bench,一个可复现的基准测试与证据门控评估框架,用于检验模型输出是否满足明确的、可图验证的证据要求。该框架基于多层临床生活方式知识图谱,整合了生物医学核心(UMLS、DrugBank、SIDER)、可计算的ADA标准护理规则,以及通过机制桥梁连接的生活方式知识与血糖实验室效应。在涵盖诊断、用药安全及对抗性生活方式冲突的100个结构化病例中,基线模型在35%的GPT-4o-mini输出和33%的GPT-4o输出中未能通过证据路径检查。证据门控机制可识别无依据的遗漏,并通过约束性修订使输出达到验证级别合规。结果表明,可计算的证据约束能将未经证实的临床疏漏显式化、可度量且可纠正,适用于聚焦糖尿病的LLM输出评估。
原文摘要 · Abstract (English)
Large language models (LLMs) can produce clinically fluent recommendations for type 2 diabetes while failing to satisfy guideline constraints or explicitly justify lifestyle-related glycemic claims. We present T2D-Bench, a reproducible benchmark and evidence-gated evaluation framework for testing whether LLM outputs satisfy explicit, graph-checkable evidence requirements. T2D-Bench is built on a multi-layer clinical-lifestyle knowledge graph that combines a biomedical spine (UMLS, DrugBank, SIDER), computable ADA Standards of Care rules, and lifestyle knowledge connected through a mechanistic bridge to glycemic laboratory effects. Across 100 structured vignettes spanning diagnosis, medication safety, and adversarial lifestyle conflicts, baseline outputs failed benchmark-defined evidence-path checks in 35% of cases for GPT-4o-mini and 33% for GPT-4o. The evidence gate detects unsupported omissions and uses constrained revision to bring outputs into verifier-level compliance with benchmark-defined evidence requirements. These results show that computable evidence constraints can make unsupported clinical omissions explicit, measurable, and correctable in diabetes-focused LLM outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。