分析大模型评分器在多语言和复杂任务中的表现与局限
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
- 用英语训练的评分器可跨语言迁移,语言能力受评估语境影响更大
- 对事实错误、文化误解和不当语言等问题识别能力弱,常漏判
- 在复杂推理题上表现差,适合关注评测可靠性的人参考
LLM-as-a-Judge 和奖励模型是大型语言模型评估中替代多选题或人工标注的常用方法,尤其在长文本评价中表现突出,广泛用于排行榜评分及强化学习对齐。然而,其在非英语提示、事实验证和挑战性问题等场景下的有效性尚未充分探索。本文全面分析自动化评估工具,发现:英语评估能力显著影响其他语言的评估表现,甚至超过语言熟练度本身;大模型常无法识别并惩罚事实错误、文化误读和不当语言;最先进的评估器在英文和韩文的复杂提示下均表现不佳,暴露出对复杂推理任务的评估局限。相关数据集和代码已公开。
原文摘要 · Abstract (English)
LLM-as-a-Judge and reward models are widely used alternatives of multiple-choice questions or human annotators for large language model (LLM) evaluation. Their efficacy shines in evaluating long-form responses, serving a critical role as evaluators of leaderboards and as proxies to align LLMs via reinforcement learning. However, despite their popularity, their effectiveness in diverse contexts, such as non-English prompts, factual verification, or challenging questions, remains unexplored. In this paper, we conduct a comprehensive analysis of automated evaluators, reporting several key findings on their behavior. First, we discover that English evaluation capabilities significantly influence language-specific evaluation capabilities, often more than the language proficiency itself, enabling evaluators trained in English to easily transfer their skills to other languages. Second, we identify critical shortcomings, where LLMs fail to detect and penalize errors, such as factual inaccuracies, cultural misrepresentations, and the presence of unwanted language. Finally, we find that state-of-the-art evaluators struggle with challenging prompts, in either English or Korean, underscoring their limitations in assessing or generating complex reasoning questions. We release the dataset and codes used.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。