统一静态与交互式评估,让大模型排名更可靠
DualEval: Joint Model-Item Calibration for Unified LLM Evaluation
- 将模型能力与题目难度、锐度联合建模,共享潜在空间
- 在4个领域用18个前沿模型验证,结果平衡且可信
- 适合需要高效、可解释评估的AI研发团队
当前大模型评估依赖两类互补但常分离的信号:静态基准的客观正确标签,以及反映开放交互的竞技场偏好数据。本文提出DualEval,一种潜变量模型-项目校准框架,将模型与评估项映射到统一空间,联合估计模型能力、题目难度与锐度。在编码、数学、通用知识及日常问答四类任务中,使用18个前沿大模型、静态基准标签及经人工偏好验证的奖励模型分数进行评估。实验表明,该框架生成了可靠且均衡的模型排名,其学习到的题目级特征可支持基准压缩(提升评估效率)与异常检测(识别污染或离群样本)。总体而言,DualEval通过联合校准统一了静态与竞技场式评估,产出可解释、可审计的评估结果,推动更高效的大模型评估流程。
原文摘要 · Abstract (English)
Current LLM evaluation relies on two complementary but often disconnected signals: static benchmarks with objective correctness labels and arena-style preference data that better reflect open-ended user interactions. We introduce DualEval, a latent model-item calibration framework that represents models and evaluation items in a shared space, jointly estimating model ability together with item difficulty and sharpness. We apply DualEval across four domains: coding, math, miscellaneous domain-knowledge tasks, and generic everyday user queries. Our evaluation uses 18 frontier LLMs, static benchmark labels, and reward-model scores validated against held-out human preferences for open-ended model responses. Empirically, our framework produces reliable and balanced model rankings, and its learned item-level profiles support downstream applications such as benchmark compression for sample-efficient evaluation and anomaly detection for contamination or outlier analysis. Overall, DualEval unifies static and arena-style evaluation through joint model-item calibration, producing model rankings and item-level diagnostics that support more sample-efficient, interpretable, and auditable evaluation pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。