arXiv:2510.17028cs.CLcs.LG2025-10AAAI被引 13

通过语义扰动提升语言模型不确定性校准,解决提示敏感问题

Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language Models

  • 用改写扰动采样语义空间,缓解提示敏感导致的校准偏差
  • 新度量可分解模型不确定性,发现部分误差源于提示形式差异
  • 适合关注模型可信度与推理一致性的研究人员

大型语言模型(LLMs)存在提示敏感现象:不同但语义等价的提示会引发显著不同的输出分布,表明模型输出中的不确定性未必反映其对提示语义的真实不确定。本文将提示敏感建模为一种泛化误差,提出通过语义改写扰动采样概念空间,可在不损失准确率的前提下提升不确定性校准效果。此外,我们引入一种新的黑箱模型不确定性分解度量,相比熵基分解,该方法考虑自然语言生成中的语义连续性。实验证明,该度量能有效量化模型不确定性中由提示敏感性带来的部分。本工作为提升提示敏感语言模型的不确定性校准提供了新思路,并揭示部分模型在输入语义一致性推理上的不足。

原文摘要 · Abstract (English)

An interesting behavior in large language models (LLMs) is prompt sensitivity. When provided with different but semantically equivalent versions of the same prompt, models may produce very different distributions of answers. This suggests that the uncertainty reflected in a model's output distribution for one prompt may not reflect the model's uncertainty about the meaning of the prompt. We model prompt sensitivity as a type of generalization error, and show that sampling across the semantic ``concept space'' with paraphrasing perturbations improves uncertainty calibration without compromising accuracy. Additionally, we introduce a new metric for uncertainty decomposition in black-box LLMs that improves upon entropy-based decomposition by modeling semantic continuities in natural language generation. We show that this decomposition metric can be used to quantify how much LLM uncertainty is attributed to prompt sensitivity. Our work introduces a new way to improve uncertainty calibration in prompt-sensitive language models, and provides evidence that some LLMs fail to exhibit consistent general reasoning about the meanings of their inputs.

语言模型不确定性提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。