arXiv:2410.13824cs.CVcs.CL2024-10ICLR被引 30

用网页界面生成训练数据,提升模型理解图文混合信息能力

Harnessing Webpage UIs for Text-Rich Visual Understanding

  • 从网页可访问性树提取结构化文本,生成多模态指令
  • 在730万样本数据集上训练,网页任务性能提升48%
  • 效果延伸至文档、图表等非网页场景,泛化能力强

文本丰富的视觉理解——即处理密集文本与图像融合环境的能力——对多模态大模型有效交互结构化环境至关重要。我们提出利用基于文本的大语言模型(LLM)从网页UI中合成通用多模态指令。尽管缺乏直接视觉输入,这些文本型LLM仍可处理网页可访问性树中的结构化文本表示。随后将这些指令与网页截图配对,用于训练多模态模型。我们构建了MultiUI数据集,包含来自100万网站的730万样本,覆盖多样化的多模态任务和界面布局。在MultiUI上训练的模型不仅在网页UI任务中表现优异,在VisualWebBench上性能提升达48%,在网页代理数据集Mind2Web上的元素识别准确率提高19.1%;更令人意外的是,其泛化能力延伸至非网页任务,如文档理解、OCR和图表解析。结果表明,网页界面数据对推动各类场景下的文本丰富视觉理解具有广泛适用性。

原文摘要 · Abstract (English)

Text-rich visual understanding-the ability to process environments where dense textual content is integrated with visuals-is crucial for multimodal large language models (MLLMs) to interact effectively with structured environments. To enhance this capability, we propose synthesizing general multimodal instructions from webpage UIs using text-based large language models (LLMs). Despite lacking direct visual input, text-based LLMs are able to process structured text representations from webpage accessibility trees. These instructions are then paired with UI screenshots to train multimodal models. We introduce MultiUI, a dataset containing 7.3 million samples from 1 million websites, covering diverse multimodal tasks and UI layouts. Models trained on MultiUI not only excel in web UI tasks-achieving up to a 48% improvement on VisualWebBench and a 19.1% boost in element accuracy on a web agent dataset Mind2Web-but also generalize surprisingly well to non-web UI tasks and even to non-UI domains, such as document understanding, OCR, and chart interpretation. These results highlight the broad applicability of web UI data for advancing text-rich visual understanding across various scenarios.

视觉理解网页数据多模态训练泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。