首个支持多模态交互的手机界面智能体评测基准
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments

- 构建包含图像、音频、视频的连续多模态输入环境
- 模型在需同步音视频的任务中性能下降超40%
- 适合研究多模态智能体与人机交互的学者
现有图形界面(GUI)智能体评测主要依赖静态截图,但真实手机操作常需处理瞬时音频和时间动态视频信号。为此,我们提出OmniGUI,首个面向多模态手机环境的逐动作级评测基准。该数据集包含29个应用的709个专家示范轨迹(共2,579个操作步骤),每步提供图像、同步音频与视频片段的连续交织输入,并标注了多模态依赖程度。由于专用多模态框架尚处初期,我们采用可原生处理交织输入的基础模型作为代理基准。实验证明,当前模型在静态视觉任务中表现良好,但在需同步时序与听觉信号的场景下,动作预测性能显著下降。消融实验揭示跨模态干扰是主要瓶颈,尤其在处理无关环境噪声时。完整数据集、评估流程与基线提示已公开。
原文摘要 · Abstract (English)
Current benchmarks for graphical user interface (GUI) agents predominantly rely on static screenshots. However, real-world smartphone interaction routinely requires agents to process transient audio cues and temporal video dynamics that are tightly coupled with the moment of action. To bridge this gap, we introduce OmniGUI, the first step-level benchmark designed to evaluate GUI agents in omni-modal smartphone environments. OmniGUI provides continuous, interleaved multimodal inputs comprising static images, synchronous audio, and video clips at every action step. The dataset encompasses 709 expert-demonstrated episodes (2,579 action steps) across 29 applications, systematically annotated with objective multimodal dependency levels. Because dedicated omni-modal GUI agent frameworks are currently in their nascent stage, we select foundational omni-modal models capable of natively processing interleaved inputs to serve as agent proxies for our initial baselines. Our empirical evaluation reveals that while current models exhibit competency on visually static tasks, their action prediction performance degrades significantly in environments requiring synchronous temporal and auditory signals. Furthermore, ablation studies isolate specific operational bottlenecks, notably cross-modal interference when processing task-irrelevant environmental noise. The complete dataset, evaluation pipeline, and baseline prompts are provided in the supplementary material. Project page: https://omni-gui.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。