大模型在安全评估中判断不一致,尤其对金融类建议难达标。
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories

- 用无参考框架测试大模型在多维度安全评估中的表现。
- 金融等受监管领域判断错误率高,暴力等明显有害内容较可靠。
- 不同模型对同一内容分歧大,语言风格也影响判断一致性。
我们在无参考设置下评估了自动化评判模型在多维度安全评估中的表现。结果表明,大语言模型在识别受监管领域(如金融)中机器生成建议的安全问题时不可靠,但在识别暴力等明显有害内容方面相对可靠。模型判断的一致性差异显著,受评估标准、内容语言及语体风格影响。不同评判模型对同一输出的分歧程度高,跨越领域、评估标准和语言。这些发现为使用大模型作为评估工具提供了新视角,并向实践者提出若干使用建议。
原文摘要 · Abstract (English)
We evaluate the consistency of automated judges in conducting a multi-dimensional safety evaluation in a reference-free setup. Our results indicate that Large Language Models are unreliable judges in identifying safety issues related to machine-generated advice in regulated domains such as finance, although they are more reliable at identifying more overt forms of unsafe/harmful content such as violence. The degree of inconsistency in a model's judgments can vary significantly by the chosen safety criteria and can be impacted by the language of the content and its linguistic style as well. Finally, there is high disagreement among different judges for the same output, across domains, safety criteria, and languages. These findings provide new insights on the practice of using LLMs as evaluators and offer several recommendations for practitioners on how to use automated judges in practical scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。