arXiv:2608.21460cs.CVcs.AI2026-08

构建首个捕捉设计细节的真人工作流数据集,提升AI设计能力。

FigmaTrace: Capturing Creative Nuances in Human Figma Design Workflows

  • 用设计阶段划分法将视频转为3469条轨迹,保留创作过程细节。
  • 在4个新环境测试中,模型表现接近Claude-Opus-5等顶尖闭源模型。
  • 适合研究人机协作设计、生成式UI建模及创意决策建模的学者。

视觉语言模型在目标检测等客观任务上表现良好,但在主观性设计任务中仍落后。主要原因在于缺乏高质量的人类工作流数据,难以捕捉专家的设计偏好与决策过程。本文提出一套由专家定义的设计技能与最佳实践分类体系,并扩展为126个开放式、主观性强、长时程的设计任务。基于此,我们构建了FigmaTrace数据集,包含超过200小时真人操作视频,通过创新的设计阶段方法转化为3469条设计轨迹。利用该数据集训练4个模型,结果表明其在4个分布外的智能代理GUI环境中表现媲美Claude-Opus-5和GPT-5.6-Sol等前沿闭源模型。消融实验显示,设计阶段驱动的视频转轨迹方法优于传统长度基方法。对最佳模型Qwen3.8-27B的定性分析进一步揭示性能提升与数据集中设计趋势的关联。数据集与最优模型已开源,供社区使用。

原文摘要 · Abstract (English)

Vision Language Models have recently shown improvements in several objective and verifiable domains such as object detection but continue to underperform on subjective and creative design tasks. A major contributor to this performance gap is the lack of high quality human workflow data that captures a diverse set of preferences and decisions that make human experts good at design tasks. In this work, we first define a unique, expert curated taxonomy of design skills and best practices which we further expand into a set of 126 open ended, subjective, long horizon tasks. Built on top of this and expert solutions, our dataset FigmaTrace contains over 200 hours of human captured video data converted into 3469 design trajectories using a novel design phase-based method. We use our dataset to train four models and show that training on FigmaTrace leads to a performance improvement comparable to frontier closed models such as \textsc{Claude-Opus-5} and \textsc{GPT-5.6-Sol} on four out of distribution agentic GUI environments. We further perform a useful ablation to attribute these performance improvements to a design phase-based video to trajectory conversion which outperforms prior length-based conversion approaches. Finally, we perform a qualitative analysis on the best performing \textsc{Qwen3.8-27B} outputs to better correlate performance improvements to FigmaTrace's trends. We open source our dataset and the best model for the community.

设计生成工作流数据多模态人类行为建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。