arXiv:2504.20294cs.AIcs.CL2025-04被引 2

构建多模态设计精修数据集,揭示生成与修改指令差异。

mrCAD: Multimodal Refinement of Computer-aided Designs

  • 通过人机协作游戏收集6082组多模态精修指令数据
  • 发现修改指令比生成指令更依赖图像与文本结合表达
  • 验证主流视觉语言模型在精修任务上表现显著下降

人类协作的关键特征是能够对已传达的概念进行迭代精修。相比之下,尽管生成式AI在内容生成方面表现出色,但在根据具体语言指令对已有输出进行精准修改方面仍存在困难。为弥合人类与机器在编辑行为上的差距,我们提出mrCAD,一个基于通信游戏的多模态指令数据集。在每轮游戏中,参与者创建计算机辅助设计(CAD),并经过多轮精修以匹配特定目标设计。仅设计师可见目标,需通过文字、绘图或两者结合的方式向制造者下达指令。mrCAD包含6,082个通信游戏、15,163次指令-执行回合,由1,092对人类玩家完成。分析发现,生成与精修指令在图文构成上存在显著差异。以mrCAD为基准测试,我们发现当前最先进的视觉语言模型在遵循生成指令方面表现优于精修指令。这些结果为分析和建模尚未被现有数据集涵盖的多模态精修语言奠定了基础。

原文摘要 · Abstract (English)

A key feature of human collaboration is the ability to iteratively refine the concepts we have communicated. In contrast, while generative AI excels at the \textit{generation} of content, it often struggles to make specific language-guided \textit{modifications} of its prior outputs. To bridge the gap between how humans and machines perform edits, we present mrCAD, a dataset of multimodal instructions in a communication game. In each game, players created computer aided designs (CADs) and refined them over several rounds to match specific target designs. Only one player, the Designer, could see the target, and they must instruct the other player, the Maker, using text, drawing, or a combination of modalities. mrCAD consists of 6,082 communication games, 15,163 instruction-execution rounds, played between 1,092 pairs of human players. We analyze the dataset and find that generation and refinement instructions differ in their composition of drawing and text. Using the mrCAD task as a benchmark, we find that state-of-the-art VLMs are better at following generation instructions than refinement instructions. These results lay a foundation for analyzing and modeling a multimodal language of refinement that is not represented in previous datasets.

多模态设计精修视觉语言模型人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。