arXiv:2605.16386cs.CV2026-05被引 2

发现多模态大模型评分有向中间靠拢的系统性偏差,影响临床判断。

Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoring

论文配图:Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoring
图 1 · 摘自论文原文
  • 分析多模态大模型在认知评估中的评分倾向,发现其倾向集中于量表中段。
  • 零样本模型在容忍度指标上表现良好(MAE 0.67,准确率92%),但存在明显端点压缩。
  • 该偏差影响关键临床决策,尤其在极值评分时,适合医疗评估部署前校准研究。

多模态大语言模型在临床场景中作为自动评分工具日益受到关注,但其在有序临床量表上的评分行为仍不清晰。我们在两个公开数据集上,使用Shulman评分标准,对三类前沿大模型家族与监督深度学习模型在钟面绘制测试(CDT)图像评分上的表现进行基准测试。全微调视觉变换模型表现最佳(MAE 0.52,±1准确率91%),而零样本大模型在容忍度一致性上仍具竞争力(GPT-5 MAE 0.67,±1准确率92%),尽管绝对误差更高。然而,逐分项分析显示,所有三类大模型均表现出显著的中心趋势效应(系统性端点压缩):预测值系统性偏向量表中段,低分端(0到1)高估,高分端(5到4)低估。该现象在临床上至关重要的极端评分上尤为突出,直接影响认知障碍筛查决策。针对性消融实验表明,无论是否提供覆盖全量程的少样本示例,或去除提示中的临床术语,该效应均未消除。研究将大模型作为评判者的偏差问题从NLP评估扩展至临床评估领域,强调在高风险筛查流程中部署前需进行校准感知评估与后处理校准。

原文摘要 · Abstract (English)

Multimodal large language models (LLMs) are increasingly explored as automated evaluators in clinical settings, yet their scoring behavior on ordinal clinical scales remains poorly understood. We benchmark three frontier LLM families against supervised deep learning models for scoring Clock Drawing Test (CDT) images on two public datasets using the Shulman rubric. While fully fine-tuned Vision Transformers achieve the best calibration (MAE 0.52, within-1 accuracy 91%), zero-shot LLMs remain competitive on tolerance-based agreement (GPT-5 MAE 0.67, within-1 accuracy 92%) despite higher absolute error. However, per-score analysis reveals that all three LLM families exhibit a pronounced central tendency effect (systematic endpoint compression): predictions are systematically compressed toward the middle of the scale, with over-prediction at the low end (score 0 to 1) and under-prediction at the high end (score 5 to 4). This effect disproportionately affects the clinically critical extremes where accurate scoring most impacts screening decisions for cognitive impairment. Targeted ablations show that neither few-shot exemplars spanning the full score range nor removing clinical terminology from the prompt eliminates the effect. Our findings extend the LLM-as-a-judge bias literature from NLP evaluation to clinical assessment, and highlight the need for calibration-aware evaluation and post-hoc calibration before deploying LLM-based raters in high-stakes screening workflows.

多模态模型临床评估评分偏差大模型审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。