新评测框架自动评估网页生成,更贴近真人体验。
Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation

- 无参考基准,自主交互采集多模态数据
- 13个前沿大模型在互动任务中仍有提升空间
- 适合评估网页生成类LLM,尤其关注用户体验
前端网页代码已成为前沿大模型发布的核心展示界面,但以人工评判为主的评测方式难以快速扩展。现有自动化评测依赖参考实现、测试用例或固定清单,难以捕捉人类在真实使用中进行的综合判断。本文提出一种全新评测范式:无需参考、自主驱动、全面推理,并通过两个工具实现。 extbf{ extit{Cookie-Bench}}包含11个领域、54个叶子节点、1000个查询任务,覆盖静态展示与交互应用,分三个难度层级和三种目标语言组,且任务描述经重写以避免提示记忆。 extbf{ extit{Cookie-Frame}}基于弗拉维尔元认知监控理论,将证据积累与判断分离为三阶段:静态感知通过被动观察形成初步印象;代理驱动交互在自主探索中持续记录屏幕视频、音频及每步截图;动态评分在完整证据链构建后,给出功能与美学的整体评价,并结构化归因失败原因。在 extit{Cookie-Bench}上, extit{Cookie-Frame}与专家评分高度一致,同时揭示13个前沿大模型在交互式网页生成中仍有显著提升空间。
原文摘要 · Abstract (English)
Front-end web code has become a core product surface for every frontier LLM release, yet evaluating these interactive applications at development speed remains costly because human-judged leaderboards like Arena do not scale. Existing automated proxies typically lean on reference implementations, test suites, or rigid checklists, and tend to miss the reasoned synthesis a human reviewer performs over a live session. We articulate a new evaluation regime that is simultaneously reference-free, autonomously driven, and holistically reasoned, and instantiate it through two artifacts. \textbf{\dataname} is an 11-domain, 54-leaf, 1,000-query WebDev benchmark spanning both static-presentation and interactive-application tasks, balanced across three difficulty tiers and three target-language groups, with briefs rewritten to resist recall from circulated prompts. \textbf{\framename}, grounded in Flavell's metacognitive monitoring, separates evidence accumulation from judgment across three stages: Static Perception forms a first impression from passive observation; Agent-Driven Interaction explores the application autonomously while capturing continuous screen video, audio, and per-step screenshots; Dynamic Scoring issues holistic functionality and aesthetics verdicts with structured failure attribution only after the evidence chain is complete. On \dataname, \framename aligns closely with expert human ratings while surfacing substantial headroom across 13 frontier LLMs on interactive web generation. \noindenthttps://anonymous.4open.science/r/Cookie-3CE/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。