arXiv:2505.19662cs.AIcs.CV2025-05中稿 · ] Changes from pre…被引 4

构建真实场景下评估智能体的基准测试,聚焦工厂与零售现场任务。

FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks

  • 基于实地拍摄的图像视频数据,设计真实工作场景任务
  • 验证多模态大模型在真实环境中的表现可行性
  • 适合关注智能体实际应用落地的研究者与工程师

本文提出FieldWorkArena,一个面向真实现场工作的智能体评估基准。随着对智能体AI需求上升,其被用于检测和记录制造及零售环境中安全风险、流程违规等关键事件。现有多数基准集中于模拟或数字环境,而本研究致力于解决真实世界评估的核心挑战。通过访谈一线工人与管理者,精心设计多样化实地任务,并采集工厂、仓库和零售场所的现场影像资料。改进评估函数以适应多模态大模型(如GPT-4o)特性,实验表明该方法可有效评估智能体性能。同时揭示了新评估方法的优势与局限性。完整数据集与评估程序已公开,可通过 https://en-documents.research.global.fujitsu.com/fieldworkarena/ 获取。

原文摘要 · Abstract (English)

This paper introduces FieldWorkArena, a benchmark for agentic AI targeting real-world field work. With the recent increase in demand for agentic AI, they are built to detect and document safety hazards, procedural violations, and other critical incidents across real-world manufacturing and retail environments. Whereas most agentic AI benchmarks focus on performance in simulated or digital environments, our work addresses the fundamental challenge of evaluating agents in the real-world. In this paper, we improve the evaluation function from previous methods to assess the performance of agentic AI in diverse real-world tasks. Our dataset comprises on-site captured images/videos in factories, warehouses and retails. Tasks were meticulously developed through interviews with site workers and managers. Evaluation results confirmed that performance evaluation considering the characteristics of Multimodal LLM (MLLM) such as GPT-4o is feasible. Furthermore, this study identifies both the effectiveness and limitations of the proposed new evaluation methodology. The complete dataset and evaluation program are publicly accessible on the website (https://en-documents.research.global.fujitsu.com/fieldworkarena/)

智能体评估真实场景多模态模型工业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。