arXiv:2506.03083cs.DScs.AI2025-06

无需标注数据,用挑战机制验证大模型评判可信度

Algorithmically Establishing Trust in Evaluators

  • 通过连续设挑战题测试评估者,不依赖标签数据
  • 经过r轮挑战,可信度达1−(1/4)^r,可准确识别不可信评估者
  • 适合低资源语言标注场景,提升大模型部署可信度

评估者(如基于大模型的评判系统)在存在共识评价方法时才可信。传统方法要么依赖参考标签数据,要么假设评估者‘知道’正确答案,二者在无标签数据时均失效。为此,我们提出‘无数据算法’,可在无需任何标注数据的情况下严格建立评估者可信度。该算法通过连续向评估者提出挑战来检验其能力。我们证明:经过r轮挑战后,若评估者真正掌握正确标签,则其被接受的概率≥1−(1/4)^r;同时能可靠识别不可信评估者。本文提供形式化证明、实证测试,并应用于低资源语言标注中对大模型作为评判者的可信度评估。本工作为低数据场景下的评估者可信性提供了科学依据,解决了大模型规模化可信部署的关键瓶颈。

原文摘要 · Abstract (English)

An evaluator, such as an LLM-as-a-judge, is trustworthy when there exists some agreed-upon way to measure its performance as a labeller. Traditional approaches either rely on testing the evaluator against references or assume that it `knows' somehow the correct labelling. Both approaches fail when references are unavailable: the former requires data, and the latter is an assumption, not evidence. To address this, we introduce the `No-Data Algorithm', which provably establishes trust in an evaluator without requiring any labelled data. Our algorithm works by successively posing challenges to said evaluator. We prove that after $r$ challenge rounds, it accepts an evaluator which knows the correct labels with probability $ \geq 1 - (1/4)^r$, and reliably flags untrustworthy ones. We present formal proofs of correctness, empirical tests, and applications to assessing trust in LLMs-as-judges for low-resource language labelling. Our work enables scientifically-grounded evaluator trust in low-data domains, addressing a critical bottleneck for scalable, trustworthy LLM deployment.

大模型评测可信评估低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。