arXiv:2501.00199cs.CLcs.AI2025-01被引 11

用GPT-4分析患者访谈,初步验证其抑郁筛查可行性

GPT-4 on Clinic Depression Assessment: An LLM-Based Pilot Study

  • 通过复杂提示+低温设置提升模型分类准确率
  • 温度低于0.2时,准确率与F1分数达最优表现
  • 适合医疗AI工具研发者参考提示工程策略

抑郁症影响全球数百万人,早期识别可降低公共健康支出并预防并发症。然而,临床诊断依赖专业人员,存在人力短缺问题。本研究探索使用GPT-4基于访谈文本进行抑郁状态二分类评估。通过对比不同提示复杂度(简单与复杂提示)及温度参数(0.0–0.3),分析其对模型性能的影响。结果显示,当温度设置在0.0–0.2时,复杂提示下模型表现最佳;温度≥0.3后,随机性增强导致性能波动,提示复杂度优势消失。表明尽管GPT-4具备潜力,但需精细调参以保证结果稳定。该预研工作揭示了提示工程与大模型间的动态关系,为未来临床AI工具开发提供参考。

原文摘要 · Abstract (English)

Depression has impacted millions of people worldwide and has become one of the most prevalent mental disorders. Early mental disorder detection can lead to cost savings for public health agencies and avoid the onset of other major comorbidities. Additionally, the shortage of specialized personnel is a critical issue because clinical depression diagnosis is highly dependent on expert professionals and is time consuming. In this study, we explore the use of GPT-4 for clinical depression assessment based on transcript analysis. We examine the model's ability to classify patient interviews into binary categories: depressed and not depressed. A comparative analysis is conducted considering prompt complexity (e.g., using both simple and complex prompts) as well as varied temperature settings to assess the impact of prompt complexity and randomness on the model's performance. Results indicate that GPT-4 exhibits considerable variability in accuracy and F1-Score across configurations, with optimal performance observed at lower temperature values (0.0-0.2) for complex prompts. However, beyond a certain threshold (temperature >= 0.3), the relationship between randomness and performance becomes unpredictable, diminishing the gains from prompt complexity. These findings suggest that, while GPT-4 shows promise for clinical assessment, the configuration of the prompts and model parameters requires careful calibration to ensure consistent results. This preliminary study contributes to understanding the dynamics between prompt engineering and large language models, offering insights for future development of AI-powered tools in clinical settings.

抑郁症筛查大模型应用提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。