发现界面元素定位评估中,语义理解常被标签复现混淆。
Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and Web

- 用词法基线对比五种编码器,检验文本嵌入是否真懂界面
- 在无标签场景下,纯文本方法表现差,嵌入优势不明显
- 建议评估时报告词法基线和标签类型分层结果
GUI定位评估常将高指令-元素嵌入相似度视为语义对齐的证据。我们在三个移动端与网页基准上发现,这一解释常被可见标签恢复所干扰。词法基线在顶级命中率上仍具竞争力,标签缺失目标下纯文本方法表现弱,且编码器的顶级命中率可由词法排名、候选池大小和标签类型预测。我们将每项操作视为同屏排序任务,对比五种现成单向量编码器与词法基线。编码器虽能补救部分词法遗漏,但可部署融合增益远小于目标感知最优方案。结果表明,嵌入式评估可能混淆标签恢复与真实语义接地。因此,嵌入式评估应报告词法基线、标签类型分层及可部署融合诊断。我们已开源分析脚本与逐步去文本化界面面板:https://github.com/qijia123/lexical-coupling-release。
原文摘要 · Abstract (English)
GUI grounding evaluations that expose UI elements as text metadata often treat high instruction-element embedding similarity as evidence of semantic grounding. Across three mobile and web benchmarks, we show that this interpretation is frequently confounded by visible-label recovery. Lexical baselines remain competitive at top-1, label-poor targets remain weak for text-only methods, and encoder top-1 hits are predictable from lexical rank, candidate-pool size, and label type. We evaluate each action as a same-screen ranking task, comparing five off-the-shelf single-vector encoders with lexical baselines. Encoders recover some lexical misses, but deployable fusion gains are much smaller than target-aware oracle gains. These findings show that embedding-based evaluations can conflate visible-label recovery with semantic GUI grounding. Embedding-based evaluations should therefore report lexical baselines, label-type stratification, and deployable-fusion diagnostics. Our released repository provides analysis scripts and detexted per-step panels: https://github.com/qijia123/lexical-coupling-release.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。