用推理痕迹提升大模型评分的可靠性,让只有标签的数据变身为可信赖资源。
Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Raters
- 通过拒绝采样法从仅含标签的数据中推断人类评分的推理过程。
- 在多个数据集上显著提升大模型与人类评分的一致性。
- 适合需要提升评分可信度的研究者和标注团队使用。
大语言模型(LLMs)被广泛用于评估任务评分,但在主观任务中其可靠性常受限,因人类判断涉及标签之外的细微推理。思维痕迹(thinking traces)虽具高信息量,但收集与整理困难。本文提出一种人机协作框架,从仅有标签的注释中推断思维痕迹,采用简单高效的拒绝采样方法实现大规模重建。这些推断出的痕迹应用于两项互补任务:(1) 微调开放的大模型评分器;(2) 为专有大模型评分器生成更清晰的标注指南。在多个数据集上,该方法显著提升了大模型与人类评分的一致性。此外,优化后的标注指南也增强了不同大模型之间的评分一致性。结果表明,大模型可作为人类思维痕迹的实用代理,使仅含标签的语料库扩展为蕴含推理痕迹的增强资源,从而提升大模型评分的可靠性。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as raters for evaluation tasks. However, their reliability is often limited for subjective tasks, when human judgments involve subtle reasoning beyond annotation labels. Thinking traces, the reasoning behind a judgment, are highly informative but challenging to collect and curate. We present a human-LLM collaborative framework to infer thinking traces from label-only annotations. The proposed framework uses a simple and effective rejection sampling method to reconstruct these traces at scale. These inferred thinking traces are applied to two complementary tasks: (1) fine-tuning open LLM raters; and (2) synthesizing clearer annotation guidelines for proprietary LLM raters. Across multiple datasets, our methods lead to significantly improved LLM-human agreement. Additionally, the refined annotation guidelines increase agreement among different LLM models. These results suggest that LLMs can serve as practical proxies for otherwise unrevealed human thinking traces, enabling label-only corpora to be extended into thinking-trace-augmented resources that enhance the reliability of LLM raters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。