arXiv:2410.15199cs.CVcs.GR2024-10

用文字控制3D模型变形,无需训练即可实现精准修改。

CLIPtortionist: Zero-shot Text-driven Deformation for Manufactured 3D Shapes

  • 基于CLIP模型构建文本驱动的优化目标,自动计算几何特征。
  • 采用新变形模型BoxDefGraph和CMA-ES算法,避免局部最优。
  • 零样本适配,适合工业设计、快速原型调整场景。

我们提出一种零样本文本驱动的3D形状变形系统,可将输入的制造类3D网格根据文本描述进行形变。该系统通过优化一个基于预训练视觉语言模型CLIP的客观函数来实现变形。我们发现,基于CLIP的客观函数存在大量虚假局部最优解;为此,我们引入一种新型变形模型BoxDefGraph,该模型可自动从输入网格中提取,专门捕捉大多数制造物体具有的对齐矩形/圆形几何特征。随后,采用CMA-ES全局优化算法最大化目标函数,其表现优于主流梯度优化器。实验表明,本方法生成结果自然且优于多个基线模型。

原文摘要 · Abstract (English)

We propose a zero-shot text-driven 3D shape deformation system that deforms an input 3D mesh of a manufactured object to fit an input text description. To do this, our system optimizes the parameters of a deformation model to maximize an objective function based on the widely used pre-trained vision language model CLIP. We find that CLIP-based objective functions exhibit many spurious local optima; to circumvent them, we parameterize deformations using a novel deformation model called BoxDefGraph which our system automatically computes from an input mesh, the BoxDefGraph is designed to capture the object aligned rectangular/circular geometry features of most manufactured objects. We then use the CMA-ES global optimization algorithm to maximize our objective, which we find to work better than popular gradient-based optimizers. We demonstrate that our approach produces appealing results and outperforms several baselines.

3D生成文本控制变形建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。