arXiv:2504.00061cs.CLcs.AI2025-04中稿 · IISE 2025 annual c…被引 1

用大模型自动采集妇科不孕病史,小模型表现更优

Evaluating the Feasibility and Accuracy of Large Language Models for Medical History-Taking in Obstetrics and Gynecology

  • 用ChatGPT-4o-mini模拟医生问诊,提取信息更完整
  • 小模型信息提取准确率92.6%,历史采集完整率达97.6%
  • 适合临床辅助问诊系统开发,需专家验证可靠性

在不孕不育等复杂敏感领域,高效医患沟通对诊断至关重要但耗时严重。本研究评估大语言模型在不孕症病史采集中的可行性和性能。构建AI对话系统,使用ChatGPT-4o和ChatGPT-4o-mini处理70例真实不孕病例,生成420份诊断病史。通过F1分数、鉴别诊断准确率(DDs Accuracy)和不孕类型判断准确率(ITJ)评估。结果显示,ChatGPT-4o-mini在信息提取准确率(F1: 0.9258 vs. 0.9029, p=0.045)和病史完整性(97.58% vs. 77.11%)上优于ChatGPT-4o;而后者在鉴别诊断准确率上略高(2.0524 vs. 2.0048),但前者在不孕类型判断准确率更高(0.6476 vs. 0.5905),一致性较低(Cronbach's α=0.562)。两模型均具备自动化病史采集可行性,其中ChatGPT-4o-mini在完整性和准确性上更优。未来需开展临床专家验证、模型微调及更大规模混合病例数据集研究。

原文摘要 · Abstract (English)

Effective physician-patient communications in pre-diagnostic environments, and most specifically in complex and sensitive medical areas such as infertility, are critical but consume a lot of time and, therefore, cause clinic workflows to become inefficient. Recent advancements in Large Language Models (LLMs) offer a potential solution for automating conversational medical history-taking and improving diagnostic accuracy. This study evaluates the feasibility and performance of LLMs in those tasks for infertility cases. An AI-driven conversational system was developed to simulate physician-patient interactions with ChatGPT-4o and ChatGPT-4o-mini. A total of 70 real-world infertility cases were processed, generating 420 diagnostic histories. Model performance was assessed using F1 score, Differential Diagnosis (DDs) Accuracy, and Accuracy of Infertility Type Judgment (ITJ). ChatGPT-4o-mini outperformed ChatGPT-4o in information extraction accuracy (F1 score: 0.9258 vs. 0.9029, p = 0.045, d = 0.244) and demonstrated higher completeness in medical history-taking (97.58% vs. 77.11%), suggesting that ChatGPT-4o-mini is more effective in extracting detailed patient information, which is critical for improving diagnostic accuracy. In contrast, ChatGPT-4o performed slightly better in differential diagnosis accuracy (2.0524 vs. 2.0048, p > 0.05). ITJ accuracy was higher in ChatGPT-4o-mini (0.6476 vs. 0.5905) but with lower consistency (Cronbach's $α$ = 0.562), suggesting variability in classification reliability. Both models demonstrated strong feasibility in automating infertility history-taking, with ChatGPT-4o-mini excelling in completeness and extraction accuracy. In future studies, expert validation for accuracy and dependability in a clinical setting, AI model fine-tuning, and larger datasets with a mix of cases of infertility have to be prioritized.

医疗AI大模型问诊系统不孕

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。