arXiv:2501.13936cs.AIcs.CL2025-01被引 3

测试大模型在医疗数值推理中的准确率,发现加校验机制后提升11%。

Evaluating Computational Accuracy of Large Language Models in Numerical Reasoning Tasks for Healthcare Applications

  • 用提示工程与校验流水线提升模型数值计算能力。
  • 整体准确率达84.10%,多步推理仍存短板。
  • 适合关注医疗AI安全性的研究者与开发者。

大型语言模型(LLMs)在医疗领域展现出强大的自然语言理解与生成能力,但在高风险场景下的数值推理能力仍不明确。本文针对临床应用中的数值推理任务,构建包含1000个真实医疗问题的数据集,涵盖剂量计算与检验结果解读等场景,评估基于GPT-3架构的优化模型表现。研究采用提示工程、事实校验流水线及正则化技术以提升准确性与泛化能力,使用精确率、召回率与F1分数进行评估。结果显示,模型整体准确率为84.10%,简单任务表现良好,但多步推理仍有困难;引入事实校验流水线使准确率提升11%,凸显验证机制的重要性。研究揭示了LLMs在医疗数值推理中的潜力,并指明未来改进方向,助力构建可信赖、可解释且情境相关的医疗AI工具。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have emerged as transformative tools in the healthcare sector, demonstrating remarkable capabilities in natural language understanding and generation. However, their proficiency in numerical reasoning, particularly in high-stakes domains like in clinical applications, remains underexplored. Numerical reasoning is critical in healthcare applications, influencing patient outcomes, treatment planning, and resource allocation. This study investigates the computational accuracy of LLMs in numerical reasoning tasks within healthcare contexts. Using a curated dataset of 1,000 numerical problems, encompassing real-world scenarios such as dosage calculations and lab result interpretations, the performance of a refined LLM based on the GPT-3 architecture was evaluated. The methodology includes prompt engineering, integration of fact-checking pipelines, and application of regularization techniques to enhance model accuracy and generalization. Key metrics such as precision, recall, and F1-score were utilized to assess the model's efficacy. The results indicate an overall accuracy of 84.10%, with improved performance in straightforward numerical tasks and challenges in multi-step reasoning. The integration of a fact-checking pipeline improved accuracy by 11%, underscoring the importance of validation mechanisms. This research highlights the potential of LLMs in healthcare numerical reasoning and identifies avenues for further refinement to support critical decision-making in clinical environments. The findings aim to contribute to the development of reliable, interpretable, and contextually relevant AI tools for healthcare.

医疗AI数值推理大模型评估事实校验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。