用偏好采样将多维评估转为可信度得分,让选模型更简单。
Sampling Preferences Yields Simple Trustworthiness Scores
- 通过模拟用户偏好,从多维评估中提取单一可信度分数。
- 在100%情况下能完全筛选候选模型,优于帕累托最优的50%上限。
- 支持用户自定义偏好权重和信心,比平均法更灵活实用。
随着大语言模型(LLMs)的兴起,人工智能模型的表现日益多维化。为此,研究者提出了多个多维评估框架以更真实地衡量LLMs性能。然而,多维评估使决策复杂化,因缺乏明确方法选择最优模型。本文提出偏好采样(preference sampling),通过考虑用户重视的多种模型特性,从多维评估结果中提取标量可信度分数。实验基于TrustLLM与DecodingTrust的多维可信度评估,表明偏好采样在100%情况下可完全缩减候选模型集合,而帕累托最优最多仅能缩减50%。同时,偏好采样对用户先验敏感,允许指定偏好权重与置信度,而平均法则无视用户知识。
原文摘要 · Abstract (English)
With the onset of large language models (LLMs), the performance of artificial intelligence (AI) models is becoming increasingly multi-dimensional. Accordingly, there have been several large, multi-dimensional evaluation frameworks put forward to evaluate LLMs. Though these frameworks are much more realistic than previous attempts which only used a single score like accuracy, multi-dimensional evaluations can complicate decision-making since there is no obvious way to select an optimal model. This work introduces preference sampling, a method to extract a scalar trustworthiness score from multi-dimensional evaluation results by considering the many characteristics of model performance which users value. We show that preference sampling improves upon alternate aggregation methods by using multi-dimensional trustworthiness evaluations of LLMs from TrustLLM and DecodingTrust. We find that preference sampling is consistently reductive, fully reducing the set of candidate models 100% of the time whereas Pareto optimality never reduces the set by more than 50%. Likewise, preference sampling is consistently sensitive to user priors-allowing users to specify the relative weighting and confidence of their preferences-whereas averaging scores is intransigent to the users' prior knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。