arXiv:2601.21767cs.CL2026-01被引 1

评测ChatGPT在医学信息抽取任务中的表现与可靠性

Evaluating ChatGPT on Medical Information Extraction Tasks: Performance, Explainability and Beyond

  • 在6个基准数据集上测试其医学信息抽取能力
  • 性能低于微调模型,但解释质量高且忠实原文
  • 适合关注可解释性与可信度的研究者参考

大型语言模型如ChatGPT在理解用户意图和生成合理回应方面表现出色。本文系统评估了ChatGPT在4种医学信息抽取(MedIE)任务上的综合能力,涵盖6个基准数据集。通过测量其性能、可解释性、置信度、忠实度和不确定性,发现:(a) ChatGPT在MedIE任务上的表现低于微调的基线模型;(b) 能提供高质量决策解释,但预测时过度自信;(c) 在多数情况下对原始文本具有高度忠实性;(d) 生成过程中的不确定性导致信息抽取结果不可靠,可能限制其在医疗场景的应用。

原文摘要 · Abstract (English)

Large Language Models (LLMs) like ChatGPT have demonstrated amazing capabilities in comprehending user intents and generate reasonable and useful responses. Beside their ability to chat, their capabilities in various natural language processing (NLP) tasks are of interest to the research community. In this paper, we focus on assessing the overall ability of ChatGPT in 4 different medical information extraction (MedIE) tasks across 6 benchmark datasets. We present the systematically analysis by measuring ChatGPT's performance, explainability, confidence, faithfulness, and uncertainty. Our experiments reveal that: (a) ChatGPT's performance scores on MedIE tasks fall behind those of the fine-tuned baseline models. (b) ChatGPT can provide high-quality explanations for its decisions, however, ChatGPT is over-confident in its predcitions. (c) ChatGPT demonstrates a high level of faithfulness to the original text in the majority of cases. (d) The uncertainty in generation causes uncertainty in information extraction results, thus may hinder its applications in MedIE tasks.

医学信息抽取大模型评估可解释性可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。