提出更贴近用户感知的LLM一致性评估方法
Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
- 基于逻辑值构建集成方法评估LLM输出一致性
- 实验显示现有方法与人类判断差异显著(n=2976)
- 强调人工评估对模型可靠性判断的重要性
大型语言模型(LLMs)容易产生幻觉且对提示扰动敏感,常导致生成文本不一致或不可靠。为缓解此类问题,研究者提出多种衡量模型响应一致性的方法,如计算重采样响应中某结果出现的概率、分析内部状态或评估响应的logits。然而,这些方法是否真正反映用户对一致性的感知尚不明确。为此,我们开展了一项用户研究(n=2,976),发现现有方法与人类对模型一致性的判断普遍不匹配。我们提出一种基于logits的集成方法来估计一致性,其性能可媲美现有最佳度量指标。结果表明,当前自动化一致性度量存在明显缺陷,应更广泛引入人工评估,以避免因度量不完善而误判模型实际表现。
原文摘要 · Abstract (English)
Large language models (LLMs) are prone to hallucinations and sensitive to prompt perturbations, often resulting in inconsistent or unreliable generated text. Different methods have been proposed to mitigate such hallucinations and fragility, one of which is to measure the consistency of LLM responses -- the model's confidence in the response or likelihood of generating a similar response when resampled. In previous work, measuring LLM response consistency often relied on calculating the probability of a response appearing within a pool of resampled responses, analyzing internal states, or evaluating logits of responses. However, it was not clear how well these approaches approximated users' perceptions of consistency of LLM responses. To find out, we performed a user study ($n=2,976$) demonstrating that current methods for measuring LLM response consistency typically do not align well with humans' perceptions of LLM consistency. We propose a logit-based ensemble method for estimating LLM consistency and show that our method matches the performance of the best-performing existing metric in estimating human ratings of LLM consistency. Our results suggest that methods for estimating LLM consistency without human evaluation are sufficiently imperfect to warrant broader use of evaluation with human input; this would avoid misjudging the adequacy of models because of the imperfections of automated consistency metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。