让大模型评分更贴近真人,无需重训练即可显著提升一致性
Aligning Black-box Language Models with Human Judgments
- 通过线性映射学习大模型输出与人类评分的关系
- 29项任务平均一致率提升超142%,仅需少量校准样本
- 适用于小模型赶超大模型,零样本/少样本皆有效
大语言模型(LLM)正被广泛用作自动化评估推荐系统、搜索引擎等主观任务的评判工具,以替代成本高、效率低的人工评价。然而,这些系统最终服务于人类,因此大模型的评判必须与人类意见高度一致,才能保证人性化设计。由于人类评价存在个体差异和偏见,实现对齐颇具挑战。本文提出一种无需重训练或微调的简单有效框架,通过学习大模型输出与人类评分之间的线性映射,在仅使用少量校准样本的情况下,使29个任务上的平均一致率提升超过142%。该方法在零样本和少样本设置下均表现良好,在6个任务中的4个超越了人与人之间的评价一致性,并使小型模型性能达到大型模型水平。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as automated judges to evaluate recommendation systems, search engines, and other subjective tasks, where relying on human evaluators can be costly, time-consuming, and unscalable. LLMs offer an efficient solution for continuous, automated evaluation. However, since the systems that are built and improved with these judgments are ultimately designed for human use, it is crucial that LLM judgments align closely with human evaluators to ensure such systems remain human-centered. On the other hand, aligning LLM judgments with human evaluators is challenging due to individual variability and biases in human judgments. We propose a simple yet effective framework to align LLM judgments with individual human evaluators or their aggregated judgments, without retraining or fine-tuning the LLM. Our approach learns a linear mapping between the LLM's outputs and human judgments, achieving over 142% average improvement in agreement across 29 tasks with only a small number of calibration examples used for training. Notably, our method works in zero-shot and few-shot settings, exceeds inter-human agreement on four out of six tasks, and enables smaller LLMs to achieve performance comparable to that of larger models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。