用真实市场验证的创业任务测试AI代理,发现最强模型仅完成30%。
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

- 基于真实创业产品流程设计任务,而非研究者假设。
- 最强AI模型仅完成约30%的任务,多数任务有部分进展。
- 适合关注AI落地能力、真实场景评估的研究者与开发者。
大语言模型和智能体的进展显著提升了复杂任务执行能力,但现有基准多依赖研究者设定的任务,无法确认这些进步是否适用于真实用户需求。我们提出「StartupBench」,一个基于市场验证的AI创业产品端到端工作流基准。通过系统分析已获广泛采用的AI产品及其流程与用户,识别出跨专业领域具有实际需求的真实任务,并将其转化为可交付成果导向的完整任务,使用细粒度评分标准捕捉复杂要求。在统一代理框架下评估代表性模型,即使最强模型也仅能成功完成约30%的任务,尽管许多任务中已有显著部分进展。进一步分析表明,复杂指令遵循和领域专业知识是主要失败原因。结果表明,大量市场验证的工作流仍超出当前通用智能体的可靠能力范围,确立StartupBench为衡量真实用户任务端到端完成度的实证基准。
原文摘要 · Abstract (English)
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。