arXiv:2509.22926cs.CLcs.AI2025-09

测试GPT-4o在三种用药管理任务中的表现,发现准确率普遍偏低。

Large language models management of medications: three performance analyses

  • 用GPT-4o完成药物剂型匹配、药物相互作用识别和处方单生成
  • 药物剂型匹配准确率49%,平均每药遗漏1.23种、虚构1.14种
  • 处方生成无错误率达65.8%,但整体用药推荐能力仍不足

目的:大型语言模型(LLMs)在某些诊断任务中表现良好,但对其在给定诊断下推荐合理用药方案的一致性研究有限。用药管理需综合药物剂型与完整医嘱信息以确保安全。本文测试了ChatGPT可用的GPT-4o在三项用药管理任务中的表现。方法:评估GPT-4o在三类任务中的表现,包括根据通用名识别可用剂型、识别药物-药物相互作用(DDI)以及根据通用名生成用药订单。每项实验均记录模型原始文本输出,并结合临床医生评价及标准LLM指标(如TF-IDF向量、归一化莱文施泰因相似度、ROUGE 1/ROUGE L F1)进行评估。结果:在药物剂型匹配任务中,GPT-4o对通用药物的匹配准确率为49%,平均每药遗漏1.23种剂型,虚构1.14种;在药物相互作用识别任务中,准确率为54.7%;在处方生成任务中,65.8%的生成语句无药物或缩写错误。结论:模型在基础用药任务中表现持续不佳,凸显需通过临床标注数据进行领域特定训练,并建立全面评估框架用于性能基准测试。

原文摘要 · Abstract (English)

Purpose: Large language models (LLMs) have proven performance for certain diagnostic tasks, however limited studies have evaluated their consistency in recommending appropriate medication regimens for a given diagnosis. Medication management is a complex task that requires synthesis of drug formulation and complete order instructions for safe use. Here, the performance of GPT 4o, an LLM available with ChatGPT, was tested for three medication management tasks. Methods: GPT-4o performance was tested using three medication tasks: identifying available formulations for a given generic drug name, identifying drug-drug interactions (DDI) for a given medication regimen, and preparing a medication order for a given generic drug name. For each experiment, the models raw text response was captured exactly as returned and evaluated using clinician evaluation in addition to standard LLM metrics, including Term Frequency-Inverse Document Frequency (TF IDF) vectors, normalized Levenshtein similarity, and Recall-Oriented Understudy for Gisting Evaluation (ROUGE 1/ROUGE L F1) between each response and its reference string. Results: For the first task of drug-formulation matching, GPT-4o had 49% accuracy for generic medications being matched to all available formulations, with an average of 1.23 omissions per medication and 1.14 hallucinations per medication. For the second task of drug-drug interaction identification, the accuracy was 54.7% for identifying the DDI pair. For the third task, GPT-4o generated order sentences containing no medication or abbreviation errors in 65.8% of cases. Conclusions: Model performance for basic medication tasks was consistently poor. This evaluation highlights the need for domain-specific training through clinician-annotated datasets and a comprehensive evaluation framework for benchmarking performance.

大模型评测用药安全GPT-4o

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。