无需模板即可从文本生成逼真且物理可信的3D手物交互模型
THOM: Generating Physically Plausible Hand-Object Meshes From Text
- 先生成手与物体的高斯分布,再通过物理优化改进交互
- 新提取方法实现顶点与高斯点显式映射,提升拓扑一致性
- 结合视觉语言模型与接触感知优化,适合虚拟现实应用
从文本生成逼真的3D手物交互(HOI)对机器人抓取和AR/VR内容创作至关重要。然而,现有方法在视觉保真度与物理合理性之间难以兼顾,因从文本生成的高斯分布中提取网格本身病态,且结果常不适于物理优化。本文提出THOM,一种无需训练的框架,可直接从文本提示生成物理合理的3D HOI网格,无需模板物体网格。THOM采用两阶段流程:首先根据文本生成手部与物体的高斯分布,随后通过物理优化细化其交互。为确保可靠交互建模,引入一种具有显式顶点-高斯映射的网格提取方法,实现拓扑感知正则化。进一步通过接触感知优化和视觉语言模型(VLM)引导的平移精修提升物理合理性。大量实验表明,THOM生成的HOI具有强文本对齐性、高视觉真实感与良好交互合理性。
原文摘要 · Abstract (English)
Generating photorealistic 3D hand-object interactions (HOIs) from text is important for applications like robotic grasping and AR/VR content creation. In practice, however, achieving both visual fidelity and physical plausibility remains difficult, as mesh extraction from text-generated Gaussians is inherently ill-posed and the resulting meshes are often unreliable for physics-based optimization. We present THOM, a training-free framework that generates physically plausible 3D HOI meshes directly from text prompts, without requiring template object meshes. THOM follows a two-stage pipeline: it first generates hand and object Gaussians guided by text, and then refines their interaction using physics-based optimization. To enable reliable interaction modeling, we introduce a mesh extraction method with an explicit vertex-to-Gaussian mapping, which enables topology-aware regularization. We further improve physical plausibility through contact-aware optimization and vision-language model (VLM)-guided translation refinement. Extensive experiments show that THOM produces high-quality HOIs with strong text alignment, visual realism, and interaction plausibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。