用e-score提升大模型输出正确性评估的可靠性与灵活性
E-Scores for (In)Correctness Assessment of Generative Model Outputs
- 引入e值替代p值,构建更稳健的错误度量方法
- 支持数据相关容忍度设定,且误差上界可控
- 适用于数学事实与属性约束等多场景评估
尽管生成模型(尤其是大语言模型)已广泛应用,但对其输出正确性的系统性评估机制仍有限。现有基于置信预测框架的方法使用p值构造响应集合,使错误概率控制在用户设定的容忍水平内。然而,这些方法易受p-hacking影响——事后选择容忍度会破坏保证。为此,本文采用e值,为生成模型输出补充e-score作为错误度量。在保持原有保证的基础上,e-score允许用户根据数据动态设定容忍度,同时上界控制大小扭曲(一种事后错误概念)。实验验证了其在数学事实性和属性约束满足性等不同正确性形式下的有效性。
原文摘要 · Abstract (English)
While generative models, especially large language models (LLMs), are ubiquitous in today's world, principled mechanisms to assess their (in)correctness are limited. Using the conformal prediction framework, previous works construct sets of LLM responses where the probability of including an incorrect response, or error, is capped at a user-defined tolerance level. However, since these methods are based on p-values, they are susceptible to p-hacking, i.e., choosing the tolerance level post-hoc can invalidate the guarantees. We therefore leverage e-values to complement generative model outputs with e-scores as measures of incorrectness. In addition to achieving the guarantees as before, e-scores further provide users with the flexibility of choosing data-dependent tolerance levels while upper bounding size distortion, a post-hoc notion of error. We experimentally demonstrate their efficacy in assessing LLM outputs under different forms of correctness: mathematical factuality and property constraints satisfaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。