arXiv:2606.11805cs.CVcs.AI2026-06

用多视角离散表示实现文本驱动的3D手物交互建模

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization

论文配图:TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization
图 1 · 摘自论文原文
  • 分阶段生成:先产多视角视觉符号,再联合优化统一网格
  • 多视角使物体距离误差降为4.92mm,穿透体积减少至0.2193cm³
  • 适合需高精度手物交互的虚拟试穿、机器人抓取等场景

文本条件的3D生成在图像和孤立物体上发展迅速,但生成手物交互网格仍具挑战:需兼顾语言语义、跨视角一致性、物体几何、手部姿态及物理合理接触。本文提出TextHOI-3D,一种分阶段框架,利用多视角生成结果作为文本视觉生成与几何恢复之间的显式接口。该方法学习固定相机下手物观测的紧凑向量量化(VQ)空间,通过CLIP条件化的视觉自回归模型从文本预测多视角视觉符号,并基于先验初始化、多视角联合优化与抗穿透精修恢复统一手物网格。设计将语义生成与几何恢复分离,同时通过离散多视角表示保持两阶段连接。在基于HO3D的数据集评估中,相比单视角方案,多视角设置使物体CD(Chamfer Distance)从17.26 mm降至4.92 mm,穿透体积由5.3721 cm³降至0.2193 cm³,同时改善手部误差与表面F-score。结果验证了多视角视觉符号作为文本驱动3D手物网格生成的有效中间表示。

原文摘要 · Abstract (English)

Text-conditioned 3D generation has progressed rapidly for images and isolated objects, but producing a hand-object mesh remains challenging: the output must preserve language semantics, cross-view consistency, object geometry, articulated hand shape, and physically plausible contact. We present TextHOI-3D, a staged framework that uses generated multi-view observations as an explicit interface between text-conditioned visual generation and geometry-aware hand-object recovery. TextHOI-3D learns a compact VQ token space for fixed-camera hand-object observations, predicts multi-view visual tokens from text with a CLIP-conditioned visual autoregressive model, and recovers a unified hand-object mesh through prior initialization, multi-view joint optimization, and anti-penetration refinement. The design separates semantic generation from geometric recovery while keeping both stages connected by a discrete multi-view representation. On HO3D-derived evaluations, the multi-view setting reduces object CD from 17.26 mm to 4.92 mm and penetration volume from 5.3721 cm^3 to 0.2193 cm^3 compared with a single-view counterpart, while improving hand errors and surface F-scores. These results support multi-view visual tokens as an effective intermediate representation for text-driven 3D hand-object mesh creation.

3D生成手物交互多视角建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。