arXiv:2502.03971cs.CVcs.HC2025-02

用RWKV架构提升高分辨率网页理解与交互推理能力

RWKV-UI: UI Understanding with Enhanced Perception and Reasoning

  • 基于RWKV架构,引入布局检测与思维链视觉提示
  • 在高分辨率网页任务中显著提升理解与多步推理性能
  • 适合需要精准网页分析与复杂交互的AI应用

现有视觉语言模型在处理包含复杂视觉、文本和交互元素的高分辨率网页时,常面临信息丢失和推理能力有限的问题。这些问题在需要网页布局理解和多步交互推理的任务中尤为突出。为此,我们提出RWKV-UI,一种基于RWKV架构的视觉语言模型,专为高分辨率用户界面图像设计。训练过程中,引入布局检测作为视觉提示,帮助模型更好理解网页结构;同时,设计基于思维链(Chain-of-Thought)机制的视觉提示,增强模型对网页内容的理解与推理能力。实验结果表明,RWKV-UI在高分辨率网页理解与交互推理任务中表现出显著性能提升。

原文摘要 · Abstract (English)

Existing Visual Language Modelsoften struggle with information loss and limited reasoning abilities when handling high-resolution web interfaces that combine complex visual, textual, and interactive elements. These challenges are particularly evident in tasks requiring webpage layout comprehension and multi-step interactive reasoning. To address these challenges, we propose RWKV-UI, a Visual Language Model based on the RWKV architecture, specifically designed to handle high-resolution UI images. During model training, we introduce layout detection as a visual prompt to help the model better understand the webpage layout structures. Additionally, we design a visual prompt based on the Chain-of-Thought(CoT) mechanism, which enhances the model's ability to understand and reason about webpage content through reasoning chains. Experimental results show that RWKV-UI demonstrates significant performance improvements in high-resolution UI understanding and interactive reasoning tasks.

视觉语言模型网页理解交互推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。