用代码生成合成界面图,让模型更懂幻灯片和UI
DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation
- 通过代码生成带标签的合成视觉内容
- 仅需少量人工标注即可提升模型性能
- 适合无障碍与人机交互研究者
让机器理解幻灯片和用户界面等结构化视觉内容对残障人士无障碍使用至关重要。但现有方法依赖人工收集和标注数据,耗时费力。为此,我们提出一种通过代码生成带目标标签的合成结构化视觉内容的方法。该方法使用户可创建自带标签的数据集,并用少量人工标注样本训练模型。我们在三项任务中验证了效果:识别视觉元素、描述视觉内容、分类内容类型,均取得性能提升。
原文摘要 · Abstract (English)
Enabling machines to understand structured visuals like slides and user interfaces is essential for making them accessible to people with disabilities. However, achieving such understanding computationally has required manual data collection and annotation, which is time-consuming and labor-intensive. To overcome this challenge, we present a method to generate synthetic, structured visuals with target labels using code generation. Our method allows people to create datasets with built-in labels and train models with a small number of human-annotated examples. We demonstrate performance improvements in three tasks for understanding slides and UIs: recognizing visual elements, describing visual content, and classifying visual content types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。