用迭代视觉提示生成精准的UI设计评论。
Visual Prompting with Iterative Refinement for Design Critique Generation
- 通过逐轮优化文本与定位框,让大模型自动产出带区域标注的设计建议。
- 在人类专家评分中,生成效果接近人工水平,差距缩小50%。
- 适用于设计评审、目标检测等多类视觉任务,通用性强。
设计反馈对用户界面(UI)设计流程至关重要,自动化生成高质量设计批评可显著提升效率。尽管现有多模态大语言模型在诸多任务中表现优异,但在生成具视觉依据的详细设计评论方面仍存在困难——这需要评论与图像中的具体区域精确对应。本文提出一种基于迭代优化的视觉提示方法,输入为UI截图和设计规范,由大模型逐步生成包含文本评论与对应边界框的设计建议。整个过程完全由大模型驱动,利用针对每一步定制的少样本示例,不断精炼输出内容。在Gemini-1.5-pro和GPT-4o上评估表明,人类专家普遍更青睐本方法生成的评论,其性能与人工水平的差距缩小了50%(以某评分指标计)。为验证方法泛化能力,我们将其应用于开放词汇的目标与属性检测任务,结果仍优于基线方法。
原文摘要 · Abstract (English)
Feedback is crucial for every design process, such as user interface (UI) design, and automating design critiques can significantly improve the efficiency of the design workflow. Although existing multimodal large language models (LLMs) excel in many tasks, they often struggle with generating high-quality design critiques -- a complex task that requires producing detailed design comments that are visually grounded in a given design's image. Building on recent advancements in iterative refinement of text output and visual prompting methods, we propose an iterative visual prompting approach for UI critique that takes an input UI screenshot and design guidelines and generates a list of design comments, along with corresponding bounding boxes that map each comment to a specific region in the screenshot. The entire process is driven completely by LLMs, which iteratively refine both the text output and bounding boxes using few-shot samples tailored for each step. We evaluated our approach using Gemini-1.5-pro and GPT-4o, and found that human experts generally preferred the design critiques generated by our pipeline over those by the baseline, with the pipeline reducing the gap from human performance by 50% for one rating metric. To assess the generalizability of our approach to other multimodal tasks, we applied our pipeline to open-vocabulary object and attribute detection, and experiments showed that our method also outperformed the baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。