arXiv:2505.24838cs.CVcs.AI2025-05NeurIPS被引 13

首个面向3D CAD长时序操作的视频数据集与模型,助力AI理解工程级交互。

VideoCAD: A Dataset and Model for Learning Long-Horizon 3D CAD UI Interactions from Video

  • 构建41万条合成视频数据,覆盖20倍更长操作时长的3D CAD交互
  • 提出VideoCADFormer模型,直接从视频学习精准的工程界面操作
  • 支持空间推理与长序列理解,适合研究多模态大模型的工程应用

计算机辅助设计(CAD)是一项耗时且复杂的任务,需要用户对复杂3D界面进行长时间、高精度的操作。尽管近年来基于AI的界面代理取得进展,但现有数据集和方法多聚焦于移动或网页应用中的短时、低复杂度任务,难以反映专业工程工具的实际需求。本文首次提出VideoCAD,一个大规模合成数据集,包含超过41,000条由自动化框架生成的标注视频,记录了人类制作的CAD设计过程中的高保真界面操作。相比现有数据集,VideoCAD在真实工程任务上的复杂度提升一个数量级,操作时长最长可达其他数据集的20倍。我们展示了两个关键下游应用:(1) 从专业3D CAD工具中学习精确任务的界面交互;(2) 构建一个视觉问答(VQA)基准,用于评估多模态大语言模型在空间推理与视频理解方面的能力。为此,我们提出VideoCADFormer,一种基于视频直接学习CAD交互的先进模型,性能显著优于现有行为克隆基线。VideoCADFormer与衍生的VQA基准揭示了当前视频驱动界面理解的关键挑战,包括精确动作定位、多模态与空间推理能力,以及长时序依赖建模的需求。

原文摘要 · Abstract (English)

Computer-Aided Design (CAD) is a time-consuming and complex process, requiring precise, long-horizon user interactions with intricate 3D interfaces. While recent advances in AI-driven user interface (UI) agents show promise, most existing datasets and methods focus on short, low-complexity tasks in mobile or web applications, failing to capture the demands of professional engineering tools. In this work, we introduce VideoCAD, the first attempt to model UI interactions for precision engineering tasks. Specifically, VideoCAD is a large-scale synthetic dataset consisting of over 41K annotated video recordings of CAD operations, generated using an automated framework for collecting high-fidelity UI action data from human-made CAD designs. Compared to existing datasets, VideoCAD offers an order-of-magnitude increase in complexity for real-world engineering UI tasks, with time horizons up to 20x longer than those in other datasets. We show two important downstream applications of VideoCAD: (1) learning UI interactions from professional 3D CAD tools for precision tasks and (2) a visual question-answering (VQA) benchmark designed to evaluate multimodal large language models (LLMs) on spatial reasoning and video understanding. To learn the UI interactions, we propose VideoCADFormer, a state-of-the-art model for learning CAD interactions directly from video, which outperforms existing behavior cloning baselines. Both VideoCADFormer and the VQA benchmark derived from VideoCAD reveal key challenges in the current state of video-based UI understanding, including the need for precise action grounding, multi-modal and spatial reasoning, and long-horizon dependencies.

3D交互视频理解CAD建模多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。