首个面向工具使用智能体的交互式多模态评测基准,推动真实场景下智能体能力评估。
EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents

- 构建包含1045个第一视角视频任务的多模态评测集,覆盖四大日常场景。
- 最佳模型在最优场景中仅达30.62%准确率,平均为19.43%,性能受限明显。
- 引入模拟用户与动态反馈机制,支持对交互能力的客观量化评估。
随着人工智能代理在开放真实环境中的应用增多,其需融合多模态感知、基于工具的多跳推理及与用户的动态交互能力。然而,现有评测基准因难以设计紧密耦合的多能力任务、模拟自然且任务受限的用户反馈,以及保障动态交互的客观评估,无法全面衡量这些能力。为此,我们提出EgoBench,首个面向工具使用智能体的交互式多模态评测基准。EgoBench包含1,045个基于第一视角视频的任务,覆盖四个日常生活场景,并配备用户-代理-工具交互评估环境。通过三阶段协同流程,每个任务强制要求视觉感知与工具增强的多跳推理共同作用。我们还开发了多智能体模拟用户,生成高保真、任务对齐的响应以评估交互能力。此外,建立确定性联合验证框架,通过过程与结果等价性实现客观评估。在EgoBench上对八种主流视频-多模态大模型(video-MLLM)的基准测试显示,性能存在严重瓶颈:最佳模型在最优场景中仅达30.62%准确率,四场景平均为19.43%。最后,通过多维度错误分析,识别出关键失败模式,揭示未来智能体发展的能力短板。
原文摘要 · Abstract (English)
As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop reasoning, and dynamic interaction with users. However, existing benchmarks fail to jointly evaluate these capabilities due to challenges in designing strictly coupled multi-capability tasks, simulating natural and task-constrained user feedback, and ensuring objective evaluation of dynamic interaction. To bridge this gap, we introduce EgoBench, the first interactive multimodal benchmark for tool-using agents. EgoBench comprises 1,045 egocentric-video-grounded tasks covering four daily scenarios, along with a user-agent-tool interactive environment for evaluation. We implement a three-stage synergistic pipeline through which each task is designed to enforce the joint application of visual perception and tool-augmented multi-hop reasoning. We additionally develop a multi-agent simulated user within EgoBench to evaluate agents' interaction capabilities, which generates high-fidelity, task-aligned responses to agents. Furthermore, we establish a deterministic joint validation framework that guarantees objective assessment through process-based and result-based equivalence. Benchmarking eight SOTA video-MLLM agents on EgoBench reveals a severe performance ceiling: the best model achieves only 30.62% accuracy in the best-performing scenario, averaging 19.43% across all four scenarios. Finally, we conduct a multi-dimensional error analysis to disentangle failure modes, exposing capability bottlenecks for advancing future AI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。