优化提示词可让大模型评分更准,且宽松判官效果更好
Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization

- 用算法自动优化提示词,比人工设计更有效
- 宽松判官反馈使提示词提升幅度更大且更稳定
- 宽松判官训练的提示词更易跨模型通用
本研究探讨了提示词设计与评判者选择在大模型作为裁判的自由文本法律问答评估中的作用。在LEXam基准上,使用ProTeGi方法结合两个裁判(Qwen3-32B、DeepSeek-V3)的反馈,对四种任务模型进行提示词优化,并测试跨裁判迁移能力。结果表明,自动优化始终优于基线,宽松型裁判反馈带来的提升更高且更一致;由宽松裁判优化的提示词能更好迁移到严格裁判。分析显示,宽松裁判提供更宽松的反馈,生成的应用范围更广,而严格裁判导致过拟合于特定裁判。研究证明,在训练数据上算法优化提示词可超越人工设计,且裁判的评价倾向直接影响提示词的泛化能力。
原文摘要 · Abstract (English)
This work explores the role of prompt design and judge selection in LLM-as-a-Judge evaluations of free text legal question answering. We examine whether automatic task prompt optimization improves over human-centered design, whether optimization effectiveness varies by judge feedback style, and whether optimized prompts transfer across judges. We systematically address these questions on the LEXam benchmark by optimizing task prompts using the ProTeGi method with feedback from two judges (Qwen3-32B, DeepSeek-V3) across four task models, and then testing cross-judge transfer. Automatic optimization consistently outperforms the baseline, with lenient judge feedback yielding higher and more consistent gains than strict judge feedback. Prompts optimized with lenient feedback transfer better to strict judges than the reverse direction. Analysis reveals that lenient judges provide permissive feedback, yielding prompts with broader applicability, whereas strict judges produce restrictive feedback, leading to judge-specific overfitting. Our findings demonstrate algorithmically optimizing prompts on training data can outperform human-centered prompt design and that judges' dispositions during optimization shape prompt generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。