arXiv:2505.08253cs.AI2025-05被引 27

用真实任务评估大模型,发现谷歌Gemini表现最佳。

Evaluating LLM Metrics Through Real-World Capabilities

  • 基于用户实际使用场景,提炼六项核心能力进行评估。
  • 现有基准对实用能力覆盖不足,效率与可解释性欠缺。
  • 聚焦写作、总结、数据处理等日常任务,适合产品开发者参考。

随着生成式AI深入日常应用,评估其性能应更贴近真实使用场景而非抽象智能。现有基准多关注代码生成或事实记忆,但用户实际使用涵盖写作辅助、摘要、引用格式、风格反馈等多种任务。本文通过大规模问卷调查和使用日志分析,识别出六项核心实用能力:摘要、技术协助、作品审阅、数据结构化、内容生成与信息检索。评估发现,现有基准在能力覆盖、效率衡量和可解释性方面存在显著空白。基于人类中心的五项实用标准(连贯性、准确性、清晰度、相关性、效率),我们筛选出最匹配真实任务的基准,并用于对比主流模型。结果显示,在这些以实用性为导向的指标下,谷歌Gemini优于OpenAI GPT、xAI Grok、Meta LLaMA、Anthropic Claude、DeepSeek及阿里通义千问。

原文摘要 · Abstract (English)

As generative AI becomes increasingly embedded in everyday workflows, it is important to evaluate its performance in ways that reflect real-world usage rather than abstract notions of intelligence. Unlike many existing benchmarks that assess general intelligence, our approach focuses on real-world utility, evaluating how well models support users in everyday tasks. While current benchmarks emphasize code generation or factual recall, users rely on AI for a much broader range of activities-from writing assistance and summarization to citation formatting and stylistic feedback. In this paper, we analyze large-scale survey data and usage logs to identify six core capabilities that represent how people commonly use Large Language Models (LLMs): Summarization, Technical Assistance, Reviewing Work, Data Structuring, Generation, and Information Retrieval. We then assess the extent to which existing benchmarks cover these capabilities, revealing significant gaps in coverage, efficiency measurement, and interpretability. Drawing on this analysis, we use human-centered criteria to identify gaps in how well current benchmarks reflect common usage that is grounded in five practical criteria: coherence, accuracy, clarity, relevance, and efficiency. For four of the six capabilities, we identify the benchmarks that best align with real-world tasks and use them to compare leading models. We find that Google Gemini outperforms other models-including OpenAI's GPT, xAI's Grok, Meta's LLaMA, Anthropic's Claude, DeepSeek, and Qwen from Alibaba-on these utility-focused metrics.

大模型评估真实任务用户体验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。