无需人工标注,自动生成动态评分标准,提升大模型自评性能。
Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge
- 不依赖人工标注,训练免代码生成细粒度评分标准。
- 在四个基准上表现媲美现有方法,微调后超越所有基线。
- 140亿参数模型优于更大商用模型,适合评估系统优化者。
LLM-as-a-Judge 是一种可扩展的人类评估替代方案,但现有基于评分标准的方法仍依赖参考答案或专家手工制定的评分标准。我们提出一种无需训练的方法,可自动生成无须人工标注的数据集特异性和实例特异性评分标准,在四个基准测试中表现与现有方法相当。此外,我们提出一种通过元评判奖励信号迭代微调评分标准生成模型的方法。微调后的生成器在成对和点对点评估中均优于所有现有基线。值得注意的是,一个经过微调的140亿参数评分标准生成器,在评分生成性能上超过更大的专有模型,验证了微调策略的有效性。
原文摘要 · Abstract (English)
LLM-as-a-Judge is a scalable alternative to human evaluation, yet existing rubric-based methods rely on human-annotated data such as reference answers or expert-crafted rubrics. We propose to automatically generate fine-grained evaluation rubrics without any human annotation. Our training-free method generates rubrics at dataset-specific and instance-specific granularities, achieving performance competitive with existing methods across four benchmarks. We further present a method that iteratively fine-tunes a rubric generator model via meta-judge reward signals. The fine-tuned generator outperforms all existing baselines in both pairwise and pointwise evaluation. Notably, a fine-tuned 14B rubric generator outperforms a much larger proprietary model at rubric generation, showing the effectiveness of our fine-tuning strategy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。