首个评估大模型糖尿病个性化决策能力的基准,覆盖真实生活场景
DM-Bench: Benchmarking LLMs for Personalized Decision Making in Diabetes Management
- 构建涵盖7类任务的个性化问答数据集,基于1.5万患者真实血糖与行为数据
- 生成36万余条情境化问题,评估模型在准确、安全、可操作性等5项指标表现
- 揭示当前大模型在糖尿病管理中表现差异大,无全能模型,适合临床辅助研发
我们提出DM-Bench,首个针对糖尿病患者日常管理中个性化决策任务的大语言模型评估基准。不同于以往通用或面向医生的医疗基准,DM-Bench聚焦于患者端AI应用原型开发中的独特挑战,涵盖7类真实世界问题:基础血糖解读、健康教育、行为关联、高级决策与长期规划。基于来自15,000名患者(含1型、2型糖尿病及前糖尿病/健康人群)为期一个月的连续血糖监测(CGM)时序数据与行为日志(如饮食、活动),我们构建了丰富数据集,并生成共计360,600条个性化、上下文相关的问答。通过准确率、依据性、安全性、清晰度与可操作性5个维度评估8个主流大模型表现,结果揭示各模型在不同任务与指标间存在显著差异,无一模型在所有维度均占优。本基准旨在提升糖尿病护理中AI方案的可靠性、安全性与实用性。
原文摘要 · Abstract (English)
We present DM-Bench, the first benchmark designed to evaluate large language model (LLM) performance across real-world decision-making tasks faced by individuals managing diabetes in their daily lives. Unlike prior health benchmarks that are either generic, clinician-facing or focused on clinical tasks (e.g., diagnosis, triage), DM-Bench introduces a comprehensive evaluation framework tailored to the unique challenges of prototyping patient-facing AI solutions in diabetes, glucose management, metabolic health and related domains. Our benchmark encompasses 7 distinct task categories, reflecting the breadth of real-world questions individuals with diabetes ask, including basic glucose interpretation, educational queries, behavioral associations, advanced decision making and long term planning. Towards this end, we compile a rich dataset comprising one month of time-series data encompassing glucose traces and metrics from continuous glucose monitors (CGMs) and behavioral logs (e.g., eating and activity patterns) from 15,000 individuals across three different diabetes populations (type 1, type 2, pre-diabetes/general health and wellness). Using this data, we generate a total of 360,600 personalized, contextual questions across the 7 tasks. We evaluate model performance on these tasks across 5 metrics: accuracy, groundedness, safety, clarity and actionability. Our analysis of 8 recent LLMs reveals substantial variability across tasks and metrics; no single model consistently outperforms others across all dimensions. By establishing this benchmark, we aim to advance the reliability, safety, effectiveness and practical utility of AI solutions in diabetes care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。