用医学批判法量化模型与权威答案的差异,评估医疗大模型可靠性。
Med-CoDE: Medical Critique based Disagreement Evaluation Framework
- 基于医学批评构建评估框架,量化模型输出与权威答案的分歧。
- 在多个医疗数据集上验证,能有效识别模型不可靠输出。
- 适合医疗AI研发者、临床AI应用审核者使用。
大型语言模型(LLMs)在医疗领域显著提升了自动化系统生成和处理人类语言的能力。然而,其在医学场景中的可靠性与准确性仍是关键挑战。现有评估方法往往缺乏鲁棒性,难以全面评估模型性能,可能带来临床风险。本文提出Med-CoDE,一种专为医疗大模型设计的争议性评估框架。该框架采用基于批评的方法,定量衡量模型生成回答与权威医学基准之间的分歧程度,同时捕捉准确性和可靠性。通过大量实验与案例研究,验证了该框架在全面、可靠评估医疗大模型质量与可信度方面的有效性。
原文摘要 · Abstract (English)
The emergence of large language models (LLMs) has significantly influenced numerous fields, including healthcare, by enhancing the capabilities of automated systems to process and generate human-like text. However, despite their advancements, the reliability and accuracy of LLMs in medical contexts remain critical concerns. Current evaluation methods often lack robustness and fail to provide a comprehensive assessment of LLM performance, leading to potential risks in clinical settings. In this work, we propose Med-CoDE, a specifically designed evaluation framework for medical LLMs to address these challenges. The framework leverages a critique-based approach to quantitatively measure the degree of disagreement between model-generated responses and established medical ground truths. This framework captures both accuracy and reliability in medical settings. The proposed evaluation framework aims to fill the existing gap in LLM assessment by offering a systematic method to evaluate the quality and trustworthiness of medical LLMs. Through extensive experiments and case studies, we illustrate the practicality of our framework in providing a comprehensive and reliable evaluation of medical LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。