arXiv:2509.23542cs.CLcs.AI2025-09中稿 · ance被引 4

研究微调大模型裁判的寿命问题,发现其对新旧生成模型和新问题泛化能力有限。

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization

  • 在不同训练与测试分布下统一评估裁判模型的未来适应性、兼容性和问题泛化能力。
  • 多数模型难以应对未来生成模型,而对旧模型兼容性较好,DPO训练提升性能。
  • 持续学习更平衡地适应新旧响应分布,现有裁判对未见问题泛化不足。

LLM作为裁判广泛用于评估生成式模型输出及对齐奖励建模。近期,使用特定数据微调裁判模型逐渐取代直接提示前沿模型,因其在较小规模下表现更优且更具鲁棒性。然而标准评估忽略了实际部署中的三大关键问题:未来适应性与向后兼容性——即微调于当前生成模型的裁判在面对未来或过去模型生成结果时的表现;以及问题泛化——裁判在测试时对未见问题的适应能力。本文在两个推理数据集上,采用三种SFT和DPO微调算法及三类骨干模型,在统一框架下系统研究上述三点。实验表明,多数模型难以适应未来生成模型,但对旧模型兼容性良好,且DPO训练持续提升性能。持续学习相比仅训练于更强或更弱响应,能更均衡地适应分布变化。此外,所有模型在从训练中见过的问题转向未见问题时均出现性能下降,说明当前裁判未能充分泛化到未知问题。这些发现为应对不断演化的生成模型提供了实用指导。

原文摘要 · Abstract (English)

The LLM-as-a-judge paradigm is widely used in both evaluating free-text model responses and reward modeling for model alignment and fine-tuning. Recently, fine-tuning judges with judge-specific data has emerged as an often preferred choice over directly prompting frontier models as judges, as the former achieves better performance with smaller model sizes while being more robust to common biases. However, the standard evaluation ignores several practical concerns of fine-tuned judges regarding their real-world deployment. In this paper, we identify and formalize three aspects that affect the shelf life of these judges: future-proofing and backward-compatibility -- how well judges fine-tuned on responses by today's generator models perform on responses by future models or past models, as well as question generalization -- how well judges generalize to unseen questions at test time. We study these three aspects under a unified framework with varying train and test distributions in two reasoning datasets, three SFT- and DPO-based fine-tuning algorithms, and three different backbone models. Experiments suggest that future-proofing is challenging for most models, while backward-compatibility is relatively easy, with DPO-trained models consistently improving performance. We further find that continual learning provides a more balanced adaptation to shifts between older and newer response distributions than training solely on stronger or weaker responses. Moreover, all models exhibit some degree of performance degradation when moving from questions seen during training to unseen ones, showing that current judges do not fully generalize to unseen questions. These findings provide insights into practical considerations for developing and deploying judge models in the face of ever-changing generators.

大模型评估裁判模型泛化能力持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。