arXiv:2512.18080cs.HCcs.AI2025-12被引 1

首个面向人类的提示到应用生成评估基准,揭示三款工具真实体验差异。

From Prompt to Product: A Human-Centered Benchmark of Agentic App Generation Systems

  • 构建人类中心评估框架,通过真实用户对比测试系统表现。
  • 3个平台生成288个应用,Firebase Studio在易用性、信任度等维度全面领先。
  • 发现视觉美观与功能可靠存在明显差距,强调任务驱动评估必要性。

能够从自然语言提示生成全栈网页应用的智能体系统(prompt-to-app)正改变软件开发方式。然而,现有评估难以兼顾视觉效果、功能正确性和用户信任,导致对各工具的真实性能缺乏清晰判断。本文提出一个以人类为中心的评估基准,对Replit、Bolt和Firebase Studio三款主流平台进行大规模对比研究。基于96个涵盖常见网页应用任务的提示,生成288个应用实例,并通过205名参与者完成1,071次质量筛选后的两两比较,评估任务易用性、视觉吸引力、感知完整性与用户信任度。结果表明:三个系统不可互换;Firebase Studio在所有人类评估维度上均表现最佳,赢得最高胜率;Bolt在视觉方面表现尚可,但在可用性和信任度上落后;Replit则在多数指标上逊于两者。研究揭示了提示到应用系统中视觉精美与功能可靠之间的显著差距,并强调交互式、任务导向的评估至关重要。我们公开了评估框架、提示集及生成产物,支持可复现的研究与未来进展。

原文摘要 · Abstract (English)

Agentic AI systems capable of generating full-stack web applications from natural language prompts ("prompt- to-app") represent a significant shift in software development. However, evaluating these systems remains challenging, as visual polish, functional correctness, and user trust are often misaligned. As a result, it is unclear how existing prompt-to-app tools compare under realistic, human-centered evaluation criteria. In this paper, we introduce a human-centered benchmark for evaluating prompt-to-app systems and conduct a large-scale comparative study of three widely used platforms: Replit, Bolt, and Firebase Studio. Using a diverse set of 96 prompts spanning common web application tasks, we generate 288 unique application artifacts. We evaluate these systems through a large-scale human-rater study involving 205 participants and 1,071 quality-filtered pairwise comparisons, assessing task-based ease of use, visual appeal, perceived completeness, and user trust. Our results show that these systems are not interchangeable: Firebase Studio consistently outperforms competing platforms across all human-evaluated dimensions, achieving the highest win rates for ease of use, trust, visual appeal, and visual appropriateness. Bolt performs competitively on visual appeal but trails Firebase on usability and trust, while Replit underperforms relative to both across most metrics. These findings highlight a persistent gap between visual polish and functional reliability in prompt-to-app systems and demonstrate the necessity of interactive, task-based evaluation. We release our benchmark framework, prompt set, and generated artifacts to support reproducible evaluation and future research in agentic application generation.

AI编程应用生成人类评估智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。