AI生成的糖尿病咨询回复在共情和可操作性上优于医生,可辅助患者理解血糖数据。
Blinded Multi-Rater Comparative Evaluation of a Large Language Model and Clinician-Authored Responses in CGM-Informed Diabetes Counseling
- 用检索增强的AI模型生成通俗易懂的血糖监测解释,不提供个性化治疗建议。
- AI回复平均质量分4.37,高于医生回复的3.58,尤其在共情和行动指导上优势明显。
- 适合用于患者教育和就诊前准备,但不能替代医生决策或独立使用。
连续血糖监测(CGM)在糖尿病管理中至关重要,但清晰且富有同理心地解释CGM数据耗时较长。目前关于基于检索的大型语言模型(LLM)在CGM咨询中的证据有限。本研究开发了一个基于检索的LLM对话代理(CA),用于CGM解读与糖尿病咨询支持,生成通俗语言回复,避免提供个体化治疗建议。从公开数据集构建了12个CGM案例。2025年10月至2026年2月期间,6位英国资深糖尿病医生每人评估2个案例,回答24个问题。采用盲法多评审者评估,每个CA生成与医生撰写的回复由3名医生独立评分,涵盖6项质量维度。共生成288条独特回复(各144条),产生864次评分。结果显示,CA回复平均质量分(4.37)显著高于医生回复(3.58),平均差异为0.782分(95% CI 0.692–0.872;P<0.001)。最大差异出现在共情(1.062,95% CI 0.948–1.177)和可操作性(0.992,95% CI 0.877–1.106)。安全警示分布相似,两组重大关切均极少见(各3/432,0.7%)。基于检索的LLM系统可能作为辅助工具,在CGM回顾、患者教育及就诊前准备中具有价值。然而,这些发现不支持其自主进行治疗决策或在无监督环境下直接应用。
原文摘要 · Abstract (English)
Continuous glucose monitoring (CGM) is central to diabetes care, but explaining CGM patterns clearly and empathetically remains time-intensive. Evidence for retrieval-grounded large language model (LLM) systems in CGM-informed counseling remains limited. To evaluate whether a retrieval-grounded LLM-based conversational agent (CA) could support patient understanding of CGM data and preparation for routine diabetes consultations. We developed a retrieval-grounded LLM-based CA for CGM interpretation and diabetes counseling support. The system generated plain-language responses while avoiding individualized therapeutic advice. Twelve CGM-informed cases were constructed from publicly available datasets. Between Oct 2025 and Feb 2026, 6 senior UK diabetes clinicians each reviewed 2 assigned cases and answered 24 questions. In a blinded multi-rater evaluation, each CA-generated and clinician-authored response was independently rated by 3 clinicians on 6 quality dimensions. Safety flags and perceived source labels were also recorded. Primary analyses used linear mixed-effects models. A total of 288 unique responses (144 CA and 144 clinician) generated 864 ratings. The CA received higher quality scores than clinician responses (mean 4.37 vs 3.58), with an estimated mean difference of 0.782 points (95% CI 0.692-0.872; P<.001). The largest differences were for empathy (1.062, 95% CI 0.948-1.177) and actionability (0.992, 95% CI 0.877-1.106). Safety flag distributions were similar, with major concerns rare in both groups (3/432, 0.7% each). Retrieval-grounded LLM systems may have value as adjunct tools for CGM review, patient education, and preconsultation preparation. However, these findings do not support autonomous therapeutic decision-making or unsupervised real-world use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。