arXiv:2512.15617cs.CLcs.AI2025-12

用大模型当裁判评估安全风险,靠多指标加权降低出错概率。

Evaluating Metrics for Safety with LLM-as-Judges

  • 设计多权重指标组合,提升大模型评估的可靠性。
  • 通过上下文敏感判断错误严重性,动态设定人工复核阈值。
  • 适合需要高安全性的医疗、核电等关键场景使用。

大语言模型在文本处理中应用日益广泛,可替代人力完成复杂或人员不足的任务。然而,当涉及医疗随访分诊或核设施权限管理等安全关键任务时,模型出错后果严重。本文主张,不应仅依赖生成框架或图结构技术声称安全性,而应聚焦评估过程中的证据类型,尤其在采用大模型作为裁判(LLM-as-Judges, LaJ)的系统中。尽管自然语言任务难以获得确定性结果,但通过加权多指标组合,可降低评估风险;结合上下文敏感度定义错误严重性,并设置置信度阈值,在多个裁判意见不一致时触发人工介入,从而提升系统整体安全性。

原文摘要 · Abstract (English)

LLMs (Large Language Models) are increasingly used in text processing pipelines to intelligently respond to a variety of inputs and generation tasks. This raises the possibility of replacing human roles that bottleneck existing information flows, either due to insufficient staff or process complexity. However, LLMs make mistakes and some processing roles are safety critical. For example, triaging post-operative care to patients based on hospital referral letters, or updating site access schedules in nuclear facilities for work crews. If we want to introduce LLMs into critical information flows that were previously performed by humans, how can we make them safe and reliable? Rather than make performative claims about augmented generation frameworks or graph-based techniques, this paper argues that the safety argument should focus on the type of evidence we get from evaluation points in LLM processes, particularly in frameworks that employ LLM-as-Judges (LaJ) evaluators. This paper argues that although we cannot get deterministic evaluations from many natural language processing tasks, by adopting a basket of weighted metrics it may be possible to lower the risk of errors within an evaluation, use context sensitivity to define error severity and design confidence thresholds that trigger human review of critical LaJ judgments when concordance across evaluators is low.

大模型评估安全关键人工复核

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。