arXiv:2503.06029cs.CLcs.LG2025-03EMNLP被引 2

首个面向中文手机场景的LLM评估基准,填补真实使用场景评测空白。

SmartBench: Is Your LLM Truly a Good Chinese Smartphone Assistant?

  • 构建5大类20项中文手机任务,覆盖日常交互场景。
  • 每项任务含50-200组高质量问答对,支持自动化评估。
  • 支持量化后在真实手机NPU上的性能测试,适合国产助手优化。

大型语言模型(LLMs)已广泛应用于日常生活,尤其通过在智能手机上本地部署,成为智能助手。然而,现有评估基准多聚焦于英文数学与编程等客观任务,难以反映中文用户在实际移动场景中对本地化LLM的真实需求。为此,我们提出SmartBench,首个专为中文手机环境设计的本地LLM评估基准。分析主流手机厂商功能后,将其划分为五大类:文本摘要、文本问答、信息抽取、内容生成与通知管理,并细化为20个具体任务。每个任务构建包含50至200个问答对的高质量数据集,涵盖日常手机交互,并开发适配任务的自动化评估指标。我们在真实手机NPU上对多个本地LLM及多模态模型进行评估,涵盖量化部署后的表现。本工作提供中文本地LLM评估的标准框架,推动该领域进一步发展与优化。代码与数据将开源于https://github.com/vivo-ai-lab/SmartBench。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have become integral to daily life, especially advancing as intelligent assistants through on-device deployment on smartphones. However, existing LLM evaluation benchmarks predominantly focus on objective tasks like mathematics and coding in English, which do not necessarily reflect the practical use cases of on-device LLMs in real-world mobile scenarios, especially for Chinese users. To address these gaps, we introduce SmartBench, the first benchmark designed to evaluate the capabilities of on-device LLMs in Chinese mobile contexts. We analyze functionalities provided by representative smartphone manufacturers and divide them into five categories: text summarization, text Q&A, information extraction, content creation, and notification management, further detailed into 20 specific tasks. For each task, we construct high-quality datasets comprising 50 to 200 question-answer pairs that reflect everyday mobile interactions, and we develop automated evaluation criteria tailored for these tasks. We conduct comprehensive evaluations of on-device LLMs and MLLMs using SmartBench and also assess their performance after quantized deployment on real smartphone NPUs. Our contributions provide a standardized framework for evaluating on-device LLMs in Chinese, promoting further development and optimization in this critical area. Code and data will be available at https://github.com/vivo-ai-lab/SmartBench.

LLM评估中文手机本地部署SmartBench

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。