arXiv:2512.17267cs.CLcs.AI2025-12被引 4

用少量人工反馈生成可解释的自动评估指标,提升大模型评价效率。

AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators

  • 从48个评估指标库中检索并组合,结合轻量反馈生成评判标准。
  • 在5个任务上使评估相关性提升33.4%,仅需不足100条反馈数据。
  • 适合快速迭代的大模型应用,尤其适用于数据稀缺场景。

评估面向用户的人工智能应用仍是核心挑战,尤其在旅行规划、临床笔记生成或对话等开放领域。黄金标准是用户反馈(如点赞/点踩)或行为信号(如留存率),但这些在原型和研究项目中往往稀缺,或速度太慢无法用于系统优化。我们提出AutoMetrics框架,在低数据条件下合成评估指标。该框架结合从自建的MetricBank(含48项指标)中检索的结果,以及基于轻量人工反馈生成的LLM-as-a-Judge评判标准,并通过回归优化使其与人类判断的相关性最大化。AutoMetrics实现从高成本测量到可解释自动指标的转变。在5个不同任务中,其与人类评分的肯德尔相关性相比纯LLM-as-a-Judge最高提升33.4%,且仅需少于100个反馈点。我们证明,AutoMetrics可作为代理奖励,效果等同于可验证奖励。完整工具包及MetricBank已开源,以加速大模型应用的自适应评估。

原文摘要 · Abstract (English)

Evaluating user-facing AI applications remains a central challenge, especially in open-ended domains such as travel planning, clinical note generation, or dialogue. The gold standard is user feedback (e.g., thumbs up/down) or behavioral signals (e.g., retention), but these are often scarce in prototypes and research projects, or too-slow to use for system optimization. We present AutoMetrics, a framework for synthesizing evaluation metrics under low-data constraints. AutoMetrics combines retrieval from MetricBank, a collection of 48 metrics we curate, with automatically generated LLM-as-a-Judge criteria informed by lightweight human feedback. These metrics are composed via regression to maximize correlation with human signal. AutoMetrics takes you from expensive measures to interpretable automatic metrics. Across 5 diverse tasks, AutoMetrics improves Kendall correlation with human ratings by up to 33.4% over LLM-as-a-Judge while requiring fewer than 100 feedback points. We show that AutoMetrics can be used as a proxy reward to equal effect as a verifiable reward. We release the full AutoMetrics toolkit and MetricBank to accelerate adaptive evaluation of LLM applications.

自动评估大模型评测低数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。