arXiv:2601.21471cs.LGmath.OC2026-01被引 5

用大模型当裁判,智能分配人工审核,高效找最优方案

Best Arm Identification with LLM Judges and Limited Human

  • 结合大模型评分与加权校正,修正偏差并生成可靠置信区间
  • 自适应聚焦可疑情境和相近选项,减少人工审核次数
  • 适合资源有限、需精准决策的场景,如模型评估与产品优化

我们研究固定置信度下的最佳臂识别(BAI),其中每个样本可获取廉价但可能有偏的代理标签(如大语言模型裁判),而真实标签只能通过人工审计获得且成本高昂。不同于经典多保真度方法,该代理标签存在臂与上下文相关的偏差,且真实标签仅可选择性获取。标准方法可能误选最优臂,而均匀审计虽准确却浪费资源。我们证明:若无偏差校正和倾向性调整,误选概率无法收敛至零(即使代理数据无限)。为此,提出一种估计器,融合代理评分与逆倾向加权残差,并构建任意时间有效的置信序列表。基于此,设计自适应选择与审计算法,将审计集中于不可靠上下文与接近的臂。理论证明,插值式奈曼规则实现近似最优的审计效率。数值实验验证了理论结果,并展示算法在实际中的优越性能。

原文摘要 · Abstract (English)

We study fixed-confidence best-arm identification (BAI) where a cheap but potentially biased proxy (e.g., LLM judge) is available for every sample, while an expensive ground-truth label can only be acquired selectively when using a human for auditing. Unlike classical multi-fidelity BAI, the proxy is biased (arm- and context-dependent) and ground truth is selectively observed. Consequently, standard multi-fidelity methods can mis-select the best arm, and uniform auditing, though accurate, wastes scarce resources and is inefficient. We prove that without bias correction and propensity adjustment, mis-selection probability may not vanish (even with unlimited proxy data). We then develop an estimator for the mean of each arm that combines proxy scores with inverse-propensity-weighted residuals and form anytime-valid confidence sequences for that estimator. Based on the estimator and confidence sequence, we propose an algorithm that adaptively selects and audits arms. The algorithm concentrates audits on unreliable contexts and close arms and we prove that a plug-in Neyman rule achieves near-oracle audit efficiency. Numerical experiments confirm the theoretical guarantees and demonstrate the superior empirical performance of the proposed algorithm.

最佳臂识别大模型裁判自适应审计多保真度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。