arXiv:2605.04083cs.LGcs.AI2026-05被引 1

让专家评价可量化:用明确规则评估模型表现

AsymmetryZero: A Framework for Operationalizing Human Expert Preferences as Semantic Evals

论文配图:AsymmetryZero: A Framework for Operationalizing Human Expert Preferences as Semantic Evals
图 1 · 摘自论文原文
  • 设计可复用的评价契约,明确定义评分维度与标准
  • 紧凑评审团成本降低95%以上,任务结果仍稳定一致
  • 适合需要可靠评估的高阶模型部署与验证场景

当前强化学习的焦点转向评估设计:构建既能作为基准又能提供明确奖励信号的评估体系。然而,许多现实任务受主观、程序性及领域特定要求约束,难以用精确匹配目标或开放式偏好判断来表达。本文提出AsymmetryZero框架,将每项任务建模为稳定的评价契约,明确标注评分内容、评判方式及标准聚合逻辑。同一契约可通过Inspect进行模型单体评估,或通过Harbor框架实现智能体评估,确保评分可比性和审计一致性。我们以Harbor开展研究,固定任务契约,在四个前沿求解器(Claude Opus 4.6、GPT-5.4、Grok-4.20、Gemini-3.1-Pro)上对比五模型前沿评审团与五模型紧凑评审团。结果显示,准则级评审一致率在75.9%至89.6%之间(严格子集一致率77.8%至92.1%),而紧凑评审团内部分歧显著更高(3:2分歧率28.7%–32.4%),远超前沿评审团(6.1%–11.5%)。验证追踪表明,紧凑评审团将单准则判断成本降至前沿的4.2%–5.6%,延迟降至21.7%–27.1%,同时任务级结果保持高度稳定。

原文摘要 · Abstract (English)

Much of the focus in RL today is on evaluation design: building meaningful evals that serve simultaneously as benchmarks and as well-defined reward signals for post-training. Yet, many real-world tasks are governed by subjective, procedural, and domain-specific requirements that are difficult to encode as exact-match targets or open-ended preference judgments frequently used in RL pipelines today. In this work, we present AsymmetryZero, a framework for operationalizing human expert preferences as semantic evals. AsymmetryZero represents each task as a stable evaluation contract that makes grading criteria explicit: what is being graded, how each criterion is judged, and how criterion-level decisions are aggregated into a task outcome. The same contract can be executed using Inspect for model-only evaluations, as well as the Harbor Framework for agentic evaluations, enabling comparable scores and shared audit artifacts across both settings. We argue that the central challenge in post-training today is the faithful encoding of expert requirements into the evaluation itself. To that end, we present a study using Harbor that holds task contracts fixed and compares a five-model frontier jury against a five-model compact jury across four frontier-class solvers (Claude Opus 4.6, GPT-5.4, Grok-4.20, Gemini-3.1-Pro). We find that criterion-level frontier-vs-compact agreement ranges from $75.9\%$ to $89.6\%$ (strict common-subset agreement: $77.8\%$ to $92.1\%$), while compact juries exhibit substantially higher internal dissent (3--2 split rate $28.7\%$--$32.4\%$) than frontier juries ($6.1\%$--$11.5\%$). Verifier traces further show that compact juries reduce per-criterion judging cost to roughly $4.2\%$--$5.6\%$ of frontier and latency to roughly $21.7\%$--$27.1\%$, even as aggregated task-level outcomes often remain comparatively stable.

评估框架人类偏好模型评测评审团

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。