评测大模型在真实医疗问答中的表现,发现GPT-4o mini最接近专家回答。
An Empirical Evaluation of Large Language Models on Consumer Health Questions
- 用真实消费者提问数据集评估5个大模型的医疗回答能力。
- GPT-4o mini在四组评估中与专家答案最匹配,Mistral-7B表现最差。
- 适合关注医疗AI落地、模型可靠性研究的读者。
本研究在MedRedQA数据集上评估了多个大语言模型(LLMs)在消费者医疗问答中的表现。该数据集包含来自AskDocs subreddit的真实用户提问及经验证专家的回答,具有非正式语言和非专业受众的特点。使用GPT-4o mini、Llama 3.1: 70B、Mistral-123B、Mistral-7B和Gemini-Flash五种模型生成回答,并采用交叉评估方式,由各模型评价自身及他者回答,以减少偏见。结果显示,GPT-4o mini在四组模型评分中与专家答案最一致,而Mistral-7B在三组评分中得分最低。研究揭示了当前大模型在真实医疗问答中的潜力与局限,为后续优化提供方向。
原文摘要 · Abstract (English)
This study evaluates the performance of several Large Language Models (LLMs) on MedRedQA, a dataset of consumer-based medical questions and answers by verified experts extracted from the AskDocs subreddit. While LLMs have shown proficiency in clinical question answering (QA) benchmarks, their effectiveness on real-world, consumer-based, medical questions remains less understood. MedRedQA presents unique challenges, such as informal language and the need for precise responses suited to non-specialist queries. To assess model performance, responses were generated using five LLMs: GPT-4o mini, Llama 3.1: 70B, Mistral-123B, Mistral-7B, and Gemini-Flash. A cross-evaluation method was used, where each model evaluated its responses as well as those of others to minimize bias. The results indicated that GPT-4o mini achieved the highest alignment with expert responses according to four out of the five models' judges, while Mistral-7B scored lowest according to three out of five models' judges. This study highlights the potential and limitations of current LLMs for consumer health medical question answering, indicating avenues for further development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。