用可接受但语义不同响应的概率评估模型可靠性
Variability Need Not Imply Error: The Case of Adequate but Semantically Distinct Responses
- 通过分类器标注响应是否合适,计算模型生成合适答案的概率
- 在不同歧义程度提示下,该指标优于传统语义熵
- 适合评估开放问答和续写任务中模型的可信度
随着语言模型广泛应用,需评估其对提示的响应可靠性(如生成内容是否准确)。现有不确定性量化工具(如置信度、熵)常用于拒绝不确定输出。例如,Kuhn 等人(2022b)认为采样响应间的语义差异表明模型“挣扎”,可能出错。我们指出,语义多样性未必代表错误,尤其在开放性任务中,多个合理但语义不同的回答是正常的。为此,我们提出 PROBAR:通过分类器标注采样响应的适切性,并估计模型生成合适回答的概率,作为实例级可靠性指标。我们在 OPT 模型上测试该指标,在两个问答数据集和英文下一个词预测任务中验证,发现 PROBAR 在不同歧义程度的提示下均优于语义熵。
原文摘要 · Abstract (English)
With the broader use of language models (LMs) comes the need to estimate their ability to respond reliably to prompts (e.g., are generated responses likely to be correct?). Uncertainty quantification tools (notions of confidence and entropy, i.a.) can be used to that end (e.g., to reject a response when the model is `uncertain'). For example, Kuhn et al. (semantic entropy; 2022b) regard semantic variation amongst sampled responses as evidence that the model `struggles' with the prompt and that the LM is likely to err. We argue that semantic variability need not imply error--this being especially intuitive in open-ended settings, where prompts elicit multiple adequate but semantically distinct responses. Hence, we propose to annotate sampled responses for their adequacy to the prompt (e.g., using a classifier) and estimate the Probability the model assigns to Adequate Responses (PROBAR), which we then regard as an indicator of the model's reliability at the instance level. We evaluate PROBAR as a measure of confidence in selective prediction with OPT models (in two QA datasets and in next-word prediction, for English) and find PROBAR to outperform semantic entropy across prompts with varying degrees of ambiguity/open-endedness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。