arXiv:2603.02029cs.AIcs.LG2026-03被引 1

用少量人工标注+大量自动评分,实现精准高效的生成模型评估

Rich Insights from Cheap Signals: Efficient Evaluations via Tensor Factorization

  • 基于张量分解融合自动评分与少量人工标注,预训练提示与模型表征
  • 在单个提示级别预测人类偏好更准确,且置信区间更紧凑
  • 适合需要细粒度评估的模型研发与无需额外人工标注的场景

突破传统评估中将异构提示性能平均化的局限,向提示级或相对同质子集的细粒度评估迈进,是诊断生成模型优劣的关键。然而,此类评估面临数据瓶颈:人工黄金标准标签成本过高,而自动化评分常与人类判断不符。为此,我们提出一种基于张量分解的新统计模型,将廉价的自动评分数据与有限的人工黄金标准标签相结合。具体而言,该方法利用自动评分对提示和生成模型的潜在表征进行预训练,再通过小规模校准集将这些预训练表征对齐至人类偏好。该方法样本效率高,对自动评分质量鲁棒,相较于标准基线,在单提示级别预测人类偏好更准确,并为关键统计参数提供紧致置信区间。我们还展示了其实际应用价值:基于提示质量构建细粒度排行榜,并仅凭自动评分估算模型性能,无需额外人工标注。

原文摘要 · Abstract (English)

Moving beyond evaluations that collapse performance across heterogeneous prompts toward fine-grained evaluation at the prompt level, or within relatively homogeneous subsets, is necessary to diagnose generative models' strengths and weaknesses. Such fine-grained evaluations, however, suffer from a data bottleneck: human gold-standard labels are too costly at this scale, while automated ratings are often misaligned with human judgment. To resolve this challenge, we propose a novel statistical model based on tensor factorization that merges cheap autorater data with a limited set of human gold-standard labels. Specifically, our approach uses autorater scores to pretrain latent representations of prompts and generative models, and then aligns those pretrained representations to human preferences using a small calibration set. This sample-efficient methodology is robust to autorater quality, more accurately predicts human preferences on a per-prompt basis than standard baselines, and provides tight confidence intervals for key statistical parameters of interest. We also showcase the practical utility of our method by constructing granular leaderboards based on prompt qualities and by estimating model performance solely from autorater scores, eliminating the need for additional human annotations.

模型评估张量分解自动评分细粒度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。