配置变动会破坏大模型的置信预测有效性,本文系统研究其影响并提出改进方法。
Conformalized Large Language Models under Configuration Shift

- 分析提示模板、解码温度等配置变化对置信预测的影响
- 发现配置变动使实际覆盖率常低于目标值,但预测集大小基本不变
- 提出基于边界校准和感知脆弱性的新策略,可有效恢复覆盖率
置信预测(CP)是一种无需分布假设的不确定性量化框架,近期被应用于大语言模型(LLMs),在数据可交换性假设下提供有限样本覆盖保证。然而,对于LLMs而言,非符合性分数通常由推理流程决定,不仅依赖数据分布,还受提示模板、解码参数和部署设置等可配置因素影响。这些配置在实践中频繁调整,却很少被视为分布漂移的来源,其对CP有效性的影响尚不明确。我们称此为“配置漂移”,并从提示模板、解码温度和权重量化三个维度进行系统研究。在涵盖9个LLMs、4个数据集和4种非符合性分数的广泛实证研究中,我们发现配置漂移会持续削弱CP有效性,常导致实际覆盖率低于目标值。相比之下,效率基本保持:有效预测集大小仍接近独立同分布基准。我们推导出覆盖率下界,将其归因于校准与测试得分分布间的差异,并使用其有限样本插值版本作为漂移严重程度的实证诊断工具。进一步表明,这些发现可转化为实用缓解策略:基于边界的再校准在少量测试样本下有效,而脆弱性感知的校准集成可在无测试数据时显著恢复覆盖率。
原文摘要 · Abstract (English)
Conformal prediction (CP) is a distribution-free framework for uncertainty quantification that has recently been adapted to large language models (LLMs), providing prediction sets with finite-sample coverage guarantees under exchangeability. Yet for LLMs, nonconformity scores are often induced by an inference pipeline, not just a fixed model, making them depend not only on the data distribution but also on configurable factors such as the prompt template, decoding parameters, and deployment setting. Since such configurations are routinely modified in practice but rarely treated as a source of shift, their impact on CP validity remains poorly understood. We call this \emph{configuration shift} and study it systematically along three axes: prompt template, decoding temperature, and weight quantization. In a broad empirical study spanning $9$ LLMs, $4$ datasets, and $4$ nonconformity scores, we find that configuration shift consistently erodes CP validity, often driving empirical coverage below the target. By contrast, efficiency is largely preserved: valid prediction sets remain close in size to the i.i.d. baseline. We derive coverage lower bounds that attribute this loss to a discrepancy between calibration and test score distributions, and use their finite-sample plug-in versions as empirical diagnostics of shift severity. We further show that these findings lead to practical mitigations: bound-inspired recalibration is effective with limited test examples, while fragility-aware calibration ensembling recovers much of the lost coverage without test data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。