首个大规模评估AI生成应用视觉质量的基准,验证工具真实表现
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools
- 通过专家两两对比,系统评估10个AI工具生成网页的视觉质量
- 覆盖300个生成网站、4000+判断,用真技能模型给出可信排名
- 开源数据与框架,适合研究者和开发者验证与改进生成工具
AI文本转应用工具宣称可在数分钟内生成高质量应用与网站,但缺乏公开基准验证其宣称。我们提出UI-Bench,首个大规模基准,通过专家两两比较,评估10个竞争性AI文本转应用工具在视觉表现上的优劣。该基准涵盖30个提示、300个生成网站及4000余次专家判断,采用基于真技能(TrueSkill)的模型进行系统排名,并提供校准置信区间。本工作建立了可复现的AI驱动网页设计评估标准。我们发布:(i) 完整提示集,(ii) 开源评估框架,(iii) 公开排行榜。生成网站的评价结果将陆续开放。访问官网查看排行榜:https://uibench.ai/leaderboard。
原文摘要 · Abstract (English)
AI text-to-app tools promise high quality applications and websites in minutes, yet no public benchmark rigorously verifies those claims. We introduce UI-Bench, the first large-scale benchmark that evaluates visual excellence across competing AI text-to-app tools through expert pairwise comparison. Spanning 10 tools, 30 prompts, 300 generated sites, and 4,000+ expert judgments, UI-Bench ranks systems with a TrueSkill-derived model that yields calibrated confidence intervals. UI-Bench establishes a reproducible standard for advancing AI-driven web design. We release (i) the complete prompt set, (ii) an open-source evaluation framework, and (iii) a public leaderboard. The generated sites rated by participants will be released soon. View the UI-Bench leaderboard at https://uibench.ai/leaderboard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。