arXiv:2509.22957cs.LG2025-09被引 7

用大模型模拟不同人群评分,提升评测结果在真实场景的可信度

Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas

  • 用大模型扮演不同背景人类进行评分,生成带有偏见的参考数据
  • 即使评分模型或数据偏差修正不完美,仍能保证评测结果准确
  • 适合关注评测可靠性、希望低成本模拟多样用户群体的研究者

随着生成式AI应用日益广泛,评测结果的外部有效性——即从实验室环境推广到真实部署场景的能力——成为关键挑战。当评测所用的人类评分样本与系统实际部署时的目标人群分布不一致时,评估结果便可能失真。本文提出一种双重稳健的估计框架,利用大模型作为裁判(LLM-as-a-judge)模拟具有特定社会人口学特征的‘人格化’评分,生成有偏差但信息丰富的参考评分。该框架将这些不完美的‘人格化’评分与受采样偏差影响的真实人类评分结合,实现统计上有效的系统质量估计。理论上证明:只要预测人类评分的模型或纠正采样偏差的重加权模型足够准确,估计结果就有效。通过设计新型人格模拟框架(PSF),系统性地操控人格评分质量与采样偏差程度,验证了本方法的鲁棒性。该工作为融合不完美人格评分与有偏差人类评分以获得可靠评测结果提供了理论基础。

原文摘要 · Abstract (English)

As Generative AI (GenAI) systems see growing adoption, a key concern involves the external validity of evaluations, or the extent to which they generalize from lab-based to real-world deployment conditions. Threats to the external validity of GenAI evaluations arise when the source sample of human raters and system outputs used to obtain a system quality estimate differs from the target distribution at deployment time. In this work, we propose a doubly-robust estimation framework designed to address this evaluation sampling bias. Key to our approach is the use of "persona" ratings produced by prompting an LLM evaluator (i.e., an LLM-as-a-judge) to behave as a human rater with specific sociodemographic characteristics. Our doubly-robust framework combines these informative yet imperfect persona ratings with human ratings obtained under evaluation sampling bias to produce statistically valid system quality estimates. In particular, we show that our approach yields valid system quality estimates when either (i) a model trained to predict human ratings using persona ratings and source data observed under sampling bias, or (ii) a reweighting model that corrects for sampling bias is of sufficient quality. We validate our framework theoretically and via a novel Persona Simulation Framework (PSF) designed to systematically manipulate persona quality and the degree of evaluation sampling bias present in source data. Our work provides a principled foundation for combining imperfect persona ratings with human ratings observed under sampling bias to obtain valid system quality estimates.

大模型评测外部有效性双重稳健

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。