轻量级概率模型在临床诊断任务上媲美大模型,揭示其推理机制差异。
Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM
- 基于概念-诊断共现统计的朴素贝叶斯评分方法
- 性能与同源预训练大模型相当,准确率接近90%
- 提供可解释基线,适合融合改进医疗AI系统
大型语言模型(LLMs)在多项选择题临床诊断基准测试中表现优异,但其性能有多少源于潜在的概率推理尚不明确。本文以MedQA数据集中的问题为例,研究诊断选项的最优选择。提出频率基础概率排序器(FBPR),通过大规模语料库中概念-诊断共现统计的平滑朴素贝叶斯对选项进行打分。当共现统计来自OLMo和Llama的预训练语料时,FBPR的性能与相应大模型相当。直接大模型推理与FBPR在答题上差异显著,重叠度仅略高于随机水平,表明两者具有互补优势。结果强调显式概率基线的价值:提供可衡量的性能参照,并为混合模型设计提供补充信号。尽管大模型性能可能依赖于非简单频率聚合的机制,但类似传统低复杂度专家系统的策略仍能解释基准测试中相当部分的表现。
原文摘要 · Abstract (English)
Large language models (LLMs) excel on multiple-choice clinical diagnosis benchmarks, yet it is unclear how much of this performance reflects underlying probabilistic reasoning. We study this through questions from MedQA, where the task is to select the most likely diagnosis. We introduce the Frequency-Based Probabilistic Ranker (FBPR), a lightweight method that scores options with a smoothed Naive Bayes over concept-diagnosis co-occurrence statistics from a large corpus. When co-occurrence statistics were sourced from the pretraining corpora for OLMo and Llama, FBPR achieves comparable performance to the corresponding LLMs pretrained on that same corpus. Direct LLM inference and FBPR largely get different questions correct, with an overlap only slightly above random chance, indicating complementary strengths of each method. These findings highlight the continued value of explicit probabilistic baselines: they provide a meaningful performance reference point and a complementary signal for potential hybridization. While the performance of LLMs seems to be driven by a mechanism other than simple frequency aggregation, we show that an approach similar to the historically grounded, low-complexity expert systems still accounts for a substantial portion of benchmark performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。