arXiv:2603.22529cs.CVcs.AI2026-03被引 2

首个融合第一人称视频与网页任务的评测基准,让AI同时理解现实与网络。

Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos

  • 用真实第一视角视频配网页任务,实现物理世界与数字世界的联动评估
  • 构建包含电商、媒体检索等多类任务的高质量视频-任务对数据集
  • 自研评分模型达84%人类一致率,可高效评估多模态智能体性能

多模态AI智能体正日益自动化复杂的真实世界工作流,涉及在线网页操作。然而现有网页智能体评测基准存在关键缺陷:仅关注网页交互与感知,缺乏与用户真实物理环境的关联。这导致无法评估关键场景,例如智能体需通过增强现实眼镜获取的第一人称视觉信息识别周围物体,并完成相关线上任务。为此,我们提出Ego2Web,首个将第一人称视频感知与网页智能体执行相融合的基准。Ego2Web通过真实第一人称视频记录与需视觉理解、网页任务规划及在线交互才能完成的网络任务配对,构建跨物理与数字世界的评测体系。我们采用自动数据生成流程结合人工验证与修正,覆盖电商、媒体检索、知识查询等多样化任务类型,形成高质量视频-任务对。为支持准确且可扩展的评估,我们还开发了新型大模型评判方法Ego2WebJudge,其与人工判断的一致性达到约84%,显著优于现有方法。在Ego2Web上对多种SOTA智能体的实验表明,其表现普遍较弱,各类任务均存在巨大提升空间。我们还进行了全面消融实验,凸显准确视频理解在任务设计中的必要性以及当前智能体的局限性。我们希望Ego2Web能成为开发真正具备跨物理与数字世界感知与行动能力的AI助手的关键资源。

原文摘要 · Abstract (English)

Multimodal AI agents are increasingly automating complex real-world workflows that involve online web execution. However, current web-agent benchmarks suffer from a critical limitation: they focus entirely on web-based interaction and perception, lacking grounding in the user's real-world physical surroundings. This limitation prevents evaluation in crucial scenarios, such as when an agent must use egocentric visual perception (e.g., via AR glasses) to recognize an object in the user's surroundings and then complete a related task online. To address this gap, we introduce Ego2Web, the first benchmark designed to bridge egocentric video perception and web agent execution. Ego2Web pairs real-world first-person video recordings with web tasks that require visual understanding, web task planning, and interaction in an online environment for successful completion. We utilize an automatic data-generation pipeline combined with human verification and refinement to curate well-constructed, high-quality video-task pairs across diverse web task types, including e-commerce, media retrieval, knowledge lookup, etc. To facilitate accurate and scalable evaluation for our benchmark, we also develop a novel LLM-as-a-Judge automatic evaluation method, Ego2WebJudge, which achieves approximately 84% agreement with human judgment, substantially higher than existing evaluation methods. Experiments with diverse SoTA agents on our Ego2Web show that their performance is weak, with substantial headroom across all task categories. We also conduct a comprehensive ablation study on task design, highlighting the necessity of accurate video understanding in the proposed task and the limitations of current agents. We hope Ego2Web can be a critical new resource for developing truly capable AI assistants that can seamlessly see, understand, and act across the physical and digital worlds.

多模态智能体第一人称视频网页任务评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。