对话式AI通过深度问诊提升症状评估准确率,优于医生。
SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment

- 设计对话式AI主动追问症状,而非被动回应用户。
- 在1.3万+参与者中,诊断准确率比医生高2.56倍(p<0.001)。
- 适合开发医疗对话系统或研究可穿戴设备与健康关联的人。
语言模型在结构化医学案例中表现优异,但对日常症状报告的评估能力仍不明确。我们通过Fitbit应用部署了SymptomAI,一个端到端患者访谈与鉴别诊断(DDx)的对话式AI系统,随机分配13,917名参与者与五个AI代理交互。其中1,228人获得临床确诊,517例由专家团队耗时250小时标注。在盲法随机对比中,SymptomAI的诊断准确率显著高于独立临床医生(OR=2.56,p<0.001)。采用主动症状追问策略的智能体表现远超用户主导的常规对话(p<0.001)。针对1,509条来自美国大众样本的对话分析验证了结果的普适性。以SymptomAI诊断为标签,分析近400种疾病对应的50多万天可穿戴设备数据,发现急性感染与生理指标变化强相关(如流感的OR>7)。尽管存在自报金标准局限,结果表明专用完整问诊优于消费级大模型的用户引导式讨论。
原文摘要 · Abstract (English)
Language models excel at diagnostic assessments on curated medical case-studies and vignettes, performing on par with, or better than, clinical professionals. However, existing studies focus on complex scenarios with rich context making it difficult to draw conclusions about how these systems perform for patients reporting symptoms in everyday life. We deployed SymptomAI, a set of conversational AI agents for end-to-end patient interviewing and differential diagnosis (DDx), via the Fitbit app in a study that randomized participants (N=13,917) to interact with five AI agents. This corpus captures diverse communication and a realistic distribution of illnesses from a real world population. A subset of 1,228 participants reported a clinician-provided diagnosis, and 517 of these were further evaluated by a panel of clinicians during over 250 hours of annotation. SymptomAI DDx were significantly more accurate (OR = 2.56, p < 0.001) than those from independent clinicians given the same dialogue in a blinded randomized comparison. Moreover, agentic strategies which conduct a dedicated symptom interview that elicit additional symptom information before providing a diagnosis, perform substantially better than baseline, user-guided conversations (p < 0.001). An auxiliary analysis on 1,509 conversations from a general US population panel validated that these results generalize beyond wearable device users. We used SymptomAI diagnoses as labels for all 13,917 participants to analyze over 500,000 days of wearable metrics across nearly 400 unique conditions. We identified strong associations between acute infections and physiological shifts (e.g., OR > 7 for influenza). While limited by self-reported ground truth, these results demonstrate the benefits of a dedicated and complete symptom interview compared to a user-guided symptom discussion, which is the default of most consumer LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。