arXiv:2602.09007cs.AIcs.CV2026-02被引 2

评测图像生成模型在GUI交互中的动态表现,发现长序列下一致性差

GEBench: Benchmarking Image Generation Models as GUI Environments

  • 构建GEBench基准,含700个真实与虚构场景的GUI交互样本
  • 提出五维评估指标GE-Score,发现多步操作中时序连贯性不足
  • 揭示图标识别、文字渲染和定位精度是关键瓶颈,适合研究生成式UI者

近期图像生成模型已能根据用户指令预测图形用户界面(GUI)的未来状态。然而,现有基准主要关注通用视觉保真度,对GUI特定情境下的状态转换与时间连贯性评估仍不足。为填补这一空白,我们提出GEBench,一个全面评估动态交互与时间连贯性的基准。GEBench包含700个精心策划的样本,涵盖五个任务类别,覆盖单步与多步交互序列,以及定位点识别,涉及真实与虚构场景。为支持系统化评估,我们提出GE-Score,一种五维指标,用于评估目标达成、交互逻辑、内容一致性、界面合理性与视觉质量。对当前模型的广泛评估表明,尽管在单步转换上表现良好,但在长序列交互中难以维持时间连贯性与空间定位。研究发现图标理解、文本渲染与定位精度是主要瓶颈。该工作为生成式GUI环境的系统评估奠定基础,并指明未来研究方向。代码已公开:https://github.com/stepfun-ai/GEBench。

原文摘要 · Abstract (English)

Recent advancements in image generation models have enabled the prediction of future Graphical User Interface (GUI) states based on user instructions. However, existing benchmarks primarily focus on general domain visual fidelity, leaving the evaluation of state transitions and temporal coherence in GUI-specific contexts underexplored. To address this gap, we introduce GEBench, a comprehensive benchmark for evaluating dynamic interaction and temporal coherence in GUI generation. GEBench comprises 700 carefully curated samples spanning five task categories, covering both single-step interactions and multi-step trajectories across real-world and fictional scenarios, as well as grounding point localization. To support systematic evaluation, we propose GE-Score, a novel five-dimensional metric that assesses Goal Achievement, Interaction Logic, Content Consistency, UI Plausibility, and Visual Quality. Extensive evaluations on current models indicate that while they perform well on single-step transitions, they struggle significantly with maintaining temporal coherence and spatial grounding over longer interaction sequences. Our findings identify icon interpretation, text rendering, and localization precision as critical bottlenecks. This work provides a foundation for systematic assessment and suggests promising directions for future research toward building high-fidelity generative GUI environments. The code is available at: https://github.com/stepfun-ai/GEBench.

GUI生成图像生成评估基准多步交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。