用信号检测理论分析大模型的判断能力,发现温度调节会同时影响敏感性和判断标准。
LLMs as Signal Detectors: Sensitivity, Bias, and the Temperature-Criterion Analogy
- 将大模型视为信号探测器,用信号检测理论分解其判断的敏感性与偏倚
- 温度升高使模型敏感性(AUC)和判断标准同时改变,打破传统类比
- 不同模型在敏感性-偏倚空间中位置各异,仅靠校准度无法区分
大型语言模型(LLMs)的校准评估常使用期望校准误差等指标,但这些指标混淆了模型区分正确与错误答案的能力(敏感性)和其自信或谨慎回应的倾向(偏倚)。信号检测理论(SDT)可分离这两者。尽管基于SDT的指标如AUROC逐渐被采用,但完整的参数化框架——不等方差模型拟合、判别标准估计、z-ROC分析——尚未应用于LLMs。本预注册研究中,我们把三个LLM当作观察者,在168,000次事实判断任务中测试温度是否类似于人类心理物理学中的收益操纵所引发的判别标准转移。关键发现是:温度不仅改变置信度,还改变了生成答案本身,导致该类比失效。结果表明,温度同时提升敏感性(AUC)并移动判别标准。所有模型均表现出不等方差证据分布(z-ROC斜率0.52–0.84),指令型模型的不对称性更显著(0.52–0.63),而基础模型(0.77–0.87)与人类识别记忆(~0.80)更接近。SDT分解揭示,不同模型在敏感性-偏倚空间中占据不同位置,仅靠校准度无法区分,说明完整参数框架提供了现有指标无法获取的诊断信息。
原文摘要 · Abstract (English)
Large language models (LLMs) are evaluated for calibration using metrics such as Expected Calibration Error that conflate two distinct components: the model's ability to discriminate correct from incorrect answers (sensitivity) and its tendency toward confident or cautious responding (bias). Signal Detection Theory (SDT) decomposes these components. While SDT-derived metrics such as AUROC are increasingly used, the full parametric framework - unequal-variance model fitting, criterion estimation, z-ROC analysis - has not been applied to LLMs as signal detectors. In this pre-registered study, we treat three LLMs as observers performing factual discrimination across 168,000 trials and test whether temperature functions as a criterion shift analogous to payoff manipulations in human psychophysics. Critically, this analogy may break down because temperature changes the generated answer itself, not only the confidence assigned to it. Our results confirm the breakdown with temperature simultaneously increasing sensitivity (AUC) and shifting criterion. All models exhibited unequal-variance evidence distributions (z-ROC slopes 0.52-0.84), with instruct models showing more extreme asymmetry (0.52-0.63) than the base model (0.77-0.87) or human recognition memory (~0.80). The SDT decomposition revealed that models occupying distinct positions in sensitivity-bias space could not be distinguished by calibration metrics alone, demonstrating that the full parametric framework provides diagnostic information unavailable from existing metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。