arXiv:2606.16802cs.AI2026-06被引 1

构建可复现的科学仪器控制代理评测基准,解决真实设备测试难问题。

LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control

论文配图:LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control
图 1 · 摘自论文原文
  • 基于网页模拟器构建多模态图形界面任务,支持浏览器直接操作。
  • 包含96个子任务覆盖从样品装载到结果检查的完整科研流程。
  • 揭示现有智能体在反馈调节与长周期任务中仍存明显短板。

当前计算机使用评测主要聚焦虚拟系统中的软件操作,而科学仪器控制需协调复杂界面并进行反馈驱动的参数调整。直接在高精度物理设备上评估代理不现实,因成本高、存在安全风险、访问受限且难以保证结果可复现。为此,我们提出LabOSBench——一个基于网页化科学仪器模拟器的多模态GUI代理挑战性评测基准。代理可通过浏览器直接操作,避免资源密集的OS虚拟化,同时支持灵活的任务配置和执行评估。LabOSBench涵盖8个仪器模拟器,构建了96个子任务,覆盖从样品装载、对准、参数调优、数据采集到结果检验的全流程。我们评估了通用视觉语言模型、专用GUI代理模型及先进代理框架在子任务与端到端层面的表现。实验表明,尽管现有代理能完成多数结构化界面任务,但在反馈驱动操作和长周期工作流执行方面仍表现不足。总体而言,LabOSBench提供了一个可复现、低成本的测试环境,推动计算机使用代理向科学仪器控制迈进。

原文摘要 · Abstract (English)

Current computer-use benchmarks primarily focus on software operation tasks in virtualized systems, whereas scientific instrumentation scenarios require coordinated control over complex interfaces, and feedback-driven parameter adjustment. However, directly evaluating agents on physical high-precision instruments is impractical due to high cost, safety risks, limited accessibility, and difficulty in ensuring reproducible evaluation. This motivates the need for a simulated yet realistic testbed that preserves the operational challenges of scientific instruments while enabling scalable and safe benchmarking. To this end, we introduce LabOSBench, a challenging benchmark for multimodal GUI agents built on a suite of web-based scientific-instrument simulators. Operating directly via a browser, LabOSBench avoids resource-heavy OS virtualization while supporting flexible task configuration and execution-based evaluation. Specifically, LabOSBench constructs 96 subtasks across eight instrument simulators, covering workflows from sample loading, alignment, parameter tuning, and data acquisition to result inspection. We evaluate general-purpose vision-language models, specialized GUI agent models, and advanced agentic frameworks at both subtask and end-to-end levels. Our experiments reveal that while existing agents can complete many structured GUI subtasks, they still struggle with feedback-driven operations and long-horizon workflow execution. Overall, LabOSBench provides a reproducible, low-cost testbed for advancing computer-using agents toward scientific-instrument control.

智能体仪器控制评测基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。