AI聊天机器人Otiz可精准解答性病问题并提供共情支持。
Performance of a large language model-Artificial Intelligence based chatbot for counseling patients with sexually transmitted infections and genital diseases
- 基于GPT-4与有限状态机设计多代理系统,实现医疗准确与情感回应。
- 对12种疾病评估中诊断准确率达4.1-4.7分(满分5),信息正确率5.0。
- 适合临床辅助咨询,尤其缓解性病专科资源紧张问题。
全球性传播感染(STIs)负担持续上升,远超专科医生承载能力。现有聊天机器人如ChatGPT未针对性病咨询优化。我们开发了专用于性病筛查与咨询的AI聊天平台Otiz,基于GPT4-0613的多代理系统架构,融合大语言模型与确定性有限自动机原理,具备上下文相关、医学准确、共情响应能力。模块包括性病通用知识、情绪识别、急性应激障碍检测及心理治疗支持,另设并行提问建议代理。评估涵盖4种性病(肛门疣、疱疹、梅毒、尿道炎/宫颈炎)和2种非性病(念珠菌病、阴茎癌),使用模拟患者语言的提示语。每条提示由两位性病专家独立评分,采用0-5分制(0为差,5为优),共60次评估。结果显示:23位专家完成30个提示的60次评估。在性病方面,诊断准确率4.1-4.7,总体准确率4.3-4.6,信息正确率5.0,可理解性4.2-4.4,共情度4.5-4.8;但相关性得分较低(2.9-3.6),提示存在冗余。非性病诊断得分显著更低(p=0.038)。观察者间一致性良好,仅12.7%的配对评分差异超过1分。结论:类似Otiz的AI对话系统可提供准确、无评判、易懂且共情的性病相关信息,有助于减轻医疗系统压力。
原文摘要 · Abstract (English)
Introduction: Global burden of sexually transmitted infections (STIs) is rising out of proportion to specialists. Current chatbots like ChatGPT are not tailored for handling STI-related concerns out of the box. We developed Otiz, an Artificial Intelligence-based (AI-based) chatbot platform designed specifically for STI detection and counseling, and assessed its performance. Methods: Otiz employs a multi-agent system architecture based on GPT4-0613, leveraging large language model (LLM) and Deterministic Finite Automaton principles to provide contextually relevant, medically accurate, and empathetic responses. Its components include modules for general STI information, emotional recognition, Acute Stress Disorder detection, and psychotherapy. A question suggestion agent operates in parallel. Four STIs (anogenital warts, herpes, syphilis, urethritis/cervicitis) and 2 non-STIs (candidiasis, penile cancer) were evaluated using prompts mimicking patient language. Each prompt was independently graded by two venereologists conversing with Otiz as patient actors on 6 criteria using Numerical Rating Scale ranging from 0 (poor) to 5 (excellent). Results: Twenty-three venereologists did 60 evaluations of 30 prompts. Across STIs, Otiz scored highly on diagnostic accuracy (4.1-4.7), overall accuracy (4.3-4.6), correctness of information (5.0), comprehensibility (4.2-4.4), and empathy (4.5-4.8). However, relevance scores were lower (2.9-3.6), suggesting some redundancy. Diagnostic scores for non-STIs were lower (p=0.038). Inter-observer agreement was strong, with differences greater than 1 point occurring in only 12.7% of paired evaluations. Conclusions: AI conversational agents like Otiz can provide accurate, correct, discrete, non-judgmental, readily accessible and easily understandable STI-related information in an empathetic manner, and can alleviate the burden on healthcare systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。