arXiv:2508.18636cs.SEcs.AI2025-08被引 1

自动评估大模型应用质量,帮用户从海量应用中快速找到好用的。

LaQual: An Automated Framework for LLM App Quality Evaluation

  • 分三步:分类标注、静态指标筛选、动态场景测评
  • 自动评分与人工一致,可筛掉66.7%~81.3%低质应用
  • 适合开发者和平台方做应用推荐与质量管控

作为软件分发的新范式,大模型应用商店迅速兴起,为内容生成、编程辅助、教育等提供多样化选择。然而,当前应用商店的排序与推荐机制主要依赖静态指标(如用户互动、收藏数),难以高效识别高质量应用。同时,现有学术研究聚焦特定垂直领域,缺乏适用于多样化大模型应用生态的通用自动化评估框架。为此,我们提出LaQual——一个自动化的大模型应用质量评估框架。该框架包含三个关键阶段:(1) 对大模型应用进行标签化与层级分类,实现精准场景映射;(2) 采用时间加权用户行为与功能能力指标进行静态评估,过滤低质量应用;(3) 动态生成场景适配的评价指标、评分标准与任务,由大模型执行综合质量评估。在主流大模型应用商店上的实验表明,LaQual的自动评分与人工判断高度一致。通过有效筛选,可将候选应用池缩减66.7%至81.3%。用户研究表明,其相比基线系统在比较效率(均值5.45 vs. 3.30)和解释信息价值(4.75 vs. 2.25)上显著更优。结果表明,LaQual为真实场景下大模型应用的高质量发现与推荐提供了可扩展、客观且以用户为中心的解决方案。

原文摘要 · Abstract (English)

Representing a new paradigm in software distribution, LLM app stores are rapidly emerging, offering users diverse choices for content generation, coding assistance, education, and more. However, current ranking and recommendation mechanisms in LLM app stores predominantly rely on static metrics, such as user interactions and favorites, making it challenging for users to efficiently identify high-quality apps. At the same time, current academic research focuses on specific vertical fields and lacks a general, automated evaluation framework applicable to the diverse LLM app ecosystem. To address the above challenges, we present LaQual, an automated framework for LLM app quality evaluation. LaQual integrates three key stages: (1) LLM app labeling and hierarchical classification for precise scenario mapping; (2) static indicator evaluation using time-weighted user engagement and functional capability indicators to filter low-quality apps; and (3) dynamic scenario-adapted evaluation, where an LLM generates scenario-specific evaluation metrics, scoring criteria, and tasks for comprehensive quality evaluation. Experiments on a mainstream LLM app store demonstrate the effectiveness of LaQual. Its automated scores show high consistency with human judgments. Through effective screening, LaQual can reduce the candidate LLM app pool by 66.7% to 81.3%. User studies further validate its significant outperformance over baseline systems, particularly in comparison efficiency (mean 5.45 vs. 3.30) and value of explanatory information (4.75 vs. 2.25). These results demonstrate that LaQual provides a scalable, objective, and user-centric solution for high-quality discovery and recommendation of LLM apps in real-world scenarios.

大模型应用质量评估自动化评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。