arXiv:2501.07931cs.AI2025-01被引 8

ChatGPT给糖尿病自我管理建议存隐患,需改进提示与外部记忆。

Advice for Diabetes Self-Management by ChatGPT Models: Challenges and Recommendations

  • 用提示工程+外部记忆增强模型医疗建议
  • 模型常不追问信息,易给出危险错误建议
  • 适合医疗AI安全研究者与临床合作者参考

由于具备高级推理、广泛上下文理解与强大问答能力,大语言模型在医疗管理研究中日益突出。尽管能应对大量医疗问题,但在提供糖尿病等慢性病的准确、实用建议方面仍面临重大挑战。我们评估了ChatGPT 3.5和4对糖尿病患者提问的回应,分析其医学知识深度及个性化、情境化建议能力。结果发现模型存在准确性差异与内在偏见,除非使用复杂提示技术,否则难以提供定制建议。此外,两个模型常在未获取必要信息的情况下直接输出建议,可能引发危险后果。这凸显了在临床环境中缺乏人类监督时,模型的实际应用效果有限。为此,我们提出基于常识的提示评估层,并采用先进的检索增强生成技术引入疾病特异性外部记忆。该方法旨在提升信息质量,降低误导风险,推动医疗领域更可靠的AI应用。研究结果将影响未来AI在医疗中的发展方向,提升整合广度与质量。

原文摘要 · Abstract (English)

Given their ability for advanced reasoning, extensive contextual understanding, and robust question-answering abilities, large language models have become prominent in healthcare management research. Despite adeptly handling a broad spectrum of healthcare inquiries, these models face significant challenges in delivering accurate and practical advice for chronic conditions such as diabetes. We evaluate the responses of ChatGPT versions 3.5 and 4 to diabetes patient queries, assessing their depth of medical knowledge and their capacity to deliver personalized, context-specific advice for diabetes self-management. Our findings reveal discrepancies in accuracy and embedded biases, emphasizing the models' limitations in providing tailored advice unless activated by sophisticated prompting techniques. Additionally, we observe that both models often provide advice without seeking necessary clarification, a practice that can result in potentially dangerous advice. This underscores the limited practical effectiveness of these models without human oversight in clinical settings. To address these issues, we propose a commonsense evaluation layer for prompt evaluation and incorporating disease-specific external memory using an advanced Retrieval Augmented Generation technique. This approach aims to improve information quality and reduce misinformation risks, contributing to more reliable AI applications in healthcare settings. Our findings seek to influence the future direction of AI in healthcare, enhancing both the scope and quality of its integration.

医疗AI大模型糖尿病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。