用图结构建模界面布局,提升企业级UI搜索精度与效率
UISearch: Graph-Based Embeddings for Multimodal Enterprise UI Screenshots Retrieval
- 将界面截图转为带属性的图,显式编码层级与空间关系
- 在2万+金融类UI上达到0.92的Top-5准确率,平均响应47.5ms
- 支持复杂组合查询,适合需要精细区分界面的设计团队
企业软件公司维护数千个跨产品和版本的用户界面,带来设计一致性、模式发现和合规检查等重大挑战。现有方法依赖视觉相似性或文本语义,缺乏对用户界面组成中基础结构特性的显式建模。我们提出一种新颖的基于图的表示方法,将界面截图转化为编码层次关系和空间排列的属性图,可泛化至文档布局、建筑图等其他结构化视觉领域。通过对比图自编码器学习保留多层级视觉、结构和语义相似性的嵌入表示。全面分析表明,该结构嵌入在区分能力上优于当前最先进的视觉编码器,显著提升了界面表示的表达力。我们构建了UISearch系统,结合结构嵌入与语义搜索,采用可组合查询语言。在20,396个金融软件界面数据集上,系统实现0.92的Top-5准确率,中位延迟47.5毫秒(P95:124毫秒),可扩展至20,000+界面。混合索引架构支持复杂查询,实现了仅靠视觉方法无法达成的细粒度界面区分。
原文摘要 · Abstract (English)
Enterprise software companies maintain thousands of user interface screens across products and versions, creating critical challenges for design consistency, pattern discovery, and compliance check. Existing approaches rely on visual similarity or text semantics, lacking explicit modeling of structural properties fundamental to user interface (UI) composition. We present a novel graph-based representation that converts UI screenshots into attributed graphs encoding hierarchical relationships and spatial arrangements, potentially generalizable to document layouts, architectural diagrams, and other structured visual domains. A contrastive graph autoencoder learns embeddings preserving multi-level similarity across visual, structural, and semantic properties. The comprehensive analysis demonstrates that our structural embeddings achieve better discriminative power than state-of-the-art Vision Encoders, representing a fundamental advance in the expressiveness of the UI representation. We implement this representation in UISearch, a multi-modal search framework that combines structural embeddings with semantic search through a composable query language. On 20,396 financial software UIs, UISearch achieves 0.92 Top-5 accuracy with 47.5ms median latency (P95: 124ms), scaling to 20,000+ screens. The hybrid indexing architecture enables complex queries and supports fine-grained UI distinction impossible with vision-only approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。