arXiv:2605.27820cs.AI2026-05被引 2

首个面向工具使用智能体的交互式多模态评测基准,推动真实场景下智能体能力评估。

EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents

论文配图:EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents
图 1 · 摘自论文原文
  • 构建包含1045个第一视角视频任务的多模态评测集,覆盖四大日常场景。
  • 最佳模型在最优场景中仅达30.62%准确率,平均为19.43%,性能受限明显。
  • 引入模拟用户与动态反馈机制,支持对交互能力的客观量化评估。

随着人工智能代理在开放真实环境中的应用增多,其需融合多模态感知、基于工具的多跳推理及与用户的动态交互能力。然而,现有评测基准因难以设计紧密耦合的多能力任务、模拟自然且任务受限的用户反馈,以及保障动态交互的客观评估,无法全面衡量这些能力。为此,我们提出EgoBench,首个面向工具使用智能体的交互式多模态评测基准。EgoBench包含1,045个基于第一视角视频的任务,覆盖四个日常生活场景,并配备用户-代理-工具交互评估环境。通过三阶段协同流程,每个任务强制要求视觉感知与工具增强的多跳推理共同作用。我们还开发了多智能体模拟用户,生成高保真、任务对齐的响应以评估交互能力。此外,建立确定性联合验证框架,通过过程与结果等价性实现客观评估。在EgoBench上对八种主流视频-多模态大模型(video-MLLM)的基准测试显示,性能存在严重瓶颈:最佳模型在最优场景中仅达30.62%准确率,四场景平均为19.43%。最后,通过多维度错误分析,识别出关键失败模式,揭示未来智能体发展的能力短板。

原文摘要 · Abstract (English)

As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop reasoning, and dynamic interaction with users. However, existing benchmarks fail to jointly evaluate these capabilities due to challenges in designing strictly coupled multi-capability tasks, simulating natural and task-constrained user feedback, and ensuring objective evaluation of dynamic interaction. To bridge this gap, we introduce EgoBench, the first interactive multimodal benchmark for tool-using agents. EgoBench comprises 1,045 egocentric-video-grounded tasks covering four daily scenarios, along with a user-agent-tool interactive environment for evaluation. We implement a three-stage synergistic pipeline through which each task is designed to enforce the joint application of visual perception and tool-augmented multi-hop reasoning. We additionally develop a multi-agent simulated user within EgoBench to evaluate agents' interaction capabilities, which generates high-fidelity, task-aligned responses to agents. Furthermore, we establish a deterministic joint validation framework that guarantees objective assessment through process-based and result-based equivalence. Benchmarking eight SOTA video-MLLM agents on EgoBench reveals a severe performance ceiling: the best model achieves only 30.62% accuracy in the best-performing scenario, averaging 19.43% across all four scenarios. Finally, we conduct a multi-dimensional error analysis to disentangle failure modes, exposing capability bottlenecks for advancing future AI agents.

多模态智能体评测基准工具使用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。