arXiv:2506.04562cs.GRcs.CV2025-06被引 1

用视觉语言模型实现无需训练的智能网格变形

Handle-based Mesh Deformation Guided By Vision Language Model

  • 通过提示工程让VLM自动选中可变形区域和控制点
  • 多视角投票降低预测不确定性,变形更贴近用户意图
  • 无需训练、全自动,适合快速3D内容创作

网格变形是3D内容操作的基础工具。尽管已有大量研究,现有方法常存在输出质量低、需大量手动调参或依赖数据密集型训练等问题。为此,我们提出一种无需训练的基于手柄的网格变形方法。核心思想是利用视觉语言模型(VLM)通过提示工程理解并操作手柄界面。首先使用锥奇异点检测识别稀疏的潜在手柄点;随后通过提示引导VLM选择与用户指令最匹配的可变形子部分及手柄;再查询选定手柄在屏幕空间中的期望变形位置。为减少VLM预测的固有不确定性,我们设计了一种新颖的多视角投票机制。在多个基准测试中,我们的方法在CLIP和GPTEval3D评分上均更贴近用户意图,同时膜能(membrane energy)衡量的形变失真较低。总体而言,该方法无需训练、高度自动化,且始终生成高质量的网格变形。

原文摘要 · Abstract (English)

Mesh deformation is a fundamental tool in 3D content manipulation. Despite extensive prior research, existing approaches often suffer from low output quality, require significant manual tuning, or depend on data-intensive training. To address these limitations, we introduce a training-free, handle-based mesh deformation method. % Our core idea is to leverage a Vision-Language Model (VLM) to interpret and manipulate a handle-based interface through prompt engineering. We begin by applying cone singularity detection to identify a sparse set of potential handles. The VLM is then prompted to select both the deformable sub-parts of the mesh and the handles that best align with user instructions. Subsequently, we query the desired deformed positions of the selected handles in screen space. To reduce uncertainty inherent in VLM predictions, we aggregate the results from multiple camera views using a novel multi-view voting scheme. % Across a suite of benchmarks, our method produces deformations that align more closely with user intent, as measured by CLIP and GPTEval3D scores, while introducing low distortion -- quantified via membrane energy. In summary, our approach is training-free, highly automated, and consistently delivers high-quality mesh deformations.

网格变形视觉语言模型提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。