arXiv:2604.26148cs.HCcs.CL2026-04ACL

测试大模型对界面动画的理解能力,发现其识动画效果不错,但理解含义差。

Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations

  • 构建300个带标注的界面动画视频数据集,评估模型对动态界面的理解能力。
  • 模型能准确识别基础运动,但在理解动画功能和语义上远不如人类。
  • 揭示运动、上下文和感知线索是影响模型表现的关键因素,适合界面智能研究者参考。

在用户界面中,动画是传达状态与反馈的核心手段,具有重要功能性而不仅限于美观。然而,当前视觉语言模型(VLMs)在界面理解上的研究主要聚焦静态截图,对其处理动态动画的能力了解有限。为此,我们构建了AniMINT,一个包含300个密集标注的界面动画视频的新数据集。系统评估了前沿VLMs在感知动画效果、识别动画目的和解释动画意义等方面的表现。结果表明,模型能够可靠地检测基本运动,但高层语义理解仍不稳定,与人类表现存在显著差距。最后,通过引入运动、上下文和感知线索(MCPC)进行探查,揭示了影响模型性能的关键瓶颈,为未来改进指明方向。

原文摘要 · Abstract (English)

AI agents operating on user interfaces must understand how interfaces communicate state and feedback to act reliably. As a core communicative modality, animations are increasingly used in modern interfaces, serving critical functional purposes beyond mere aesthetics. Thus, understanding UI animation is essential for comprehensive interface interpretation. However, recent studies of Vision Language Models (VLMs) for UI understanding have focused primarily on static screenshots, leaving it unclear how well these models handle dynamic UI animations. To address this gap, we created AniMINT, a novel dataset of 300 densely annotated UI animation videos. We systematically evaluate state-of-the-art VLMs on UI animation understanding, including their abilities to perceive the animation effects, identify animation purposes, and interpret animation meaning. Our results show that VLMs can reliably detect primitive motion. However, their high-level animation interpretation remains inconsistent, with substantial gaps relative to human performance. Finally, we use Motion, Context, and Perceptual Cues (MCPC) to probe factors affecting VLM performance, revealing key bottlenecks and directions for future improvement.

界面理解视觉语言模型动画分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。