arXiv:2603.17172cs.LG2026-03被引 1

通过噪声干预校准大模型判官,判断其可靠性。

Noise-Response Calibration: A Causal Intervention Protocol for LLM-Judges

  • 用噪声强度变化测试模型表现,看是否随干扰下降。
  • 文本模型表现随噪声明显变差,表格数据则不敏感。
  • 模型在抗噪能力差的数据上反而表现更差,提示需警惕。

大语言模型(LLMs)被广泛用作自动判官和合成标签生成器,尤其在低标注数据场景中。然而这些系统具有随机性且常过度自信,当缺乏外部真实标签时,部署决策困难。本文提出一种基于可控输入干预的实用校准协议:若噪声强度增加,任务性能应呈现统计显著的下降趋势。我们通过重复试验中的斜率检验实现这一思路,对表格数据使用信噪比(SNR)扰动,对文本数据使用词汇扰动。在多个UCI表格基准和四个文本分类数据集上的实验显示明显的模态差异:文本判官表现随噪声显著下降,而多数表格数据集即使在显著降低信噪比的情况下也未出现统计显著的性能退化。有趣的是,模型在对噪声干预不敏感的数据集上表现更差。本文提供了一套可复现的、适用于分布偏移下的鲁棒性校准方法与报告规范。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as automated judges and synthetic labelers, especially in low-label settings. Yet these systems are stochastic and often overconfident, which makes deployment decisions difficult when external ground truth is limited. We propose a practical calibration protocol based on controlled input interventions: if noise severity increases, task performance should exhibit a statistically significant deterioration trend. We operationalize this with a slope-based hypothesis test over repeated trials, using signal-to-noise-ratio (SNR) perturbations for tabular data and lexical perturbations for text data. Across UCI tabular benchmarks and four text classification datasets, we find clear modality-dependent behavior. Our results reveal a modality gap: while text-based judges degrade predictably, the majority of tabular datasets show a lack of statistically significant performance deterioration even under significant signal-to-noise reduction. Interestingly we find that model performance is lower on datasets that are insensitive to noise interventions. We present a reproducible methodology and reporting protocol for robust LLM-judge calibration under distribution shift.

大模型评估噪声干预模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。