arXiv:2506.13023cs.AIcs.LG2025-06被引 5

为大模型应用提供真实场景下的可落地评估方案

A Practical Guide for Evaluating LLMs and LLM-Reliant Systems

  • 构建真实数据集与针对性评估指标
  • 整合开发部署流程的评估方法论
  • 适合关注实际落地的大模型开发者

生成式AI的进展推动了大量依赖大语言模型(LLMs)的应用。然而,这些系统在真实场景中的有效评估面临独特挑战,传统合成基准和常用度量无法充分应对。本文提出一个实用评估框架,指导如何主动构建代表性数据集、选择有意义的评估指标,并采用与实际开发部署流程兼容的评估方法,确保系统满足真实需求和用户期望。

原文摘要 · Abstract (English)

Recent advances in generative AI have led to remarkable interest in using systems that rely on large language models (LLMs) for practical applications. However, meaningful evaluation of these systems in real-world scenarios comes with a distinct set of challenges, which are not well-addressed by synthetic benchmarks and de-facto metrics that are often seen in the literature. We present a practical evaluation framework which outlines how to proactively curate representative datasets, select meaningful evaluation metrics, and employ meaningful evaluation methodologies that integrate well with practical development and deployment of LLM-reliant systems that must adhere to real-world requirements and meet user-facing needs.

大模型评估实证研究系统部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。