用文字指导生成和优化手物交互的3D形状,无需预设模板。
TIGeR: Text-Instructed Generation and Refinement for Template-Free Hand-Object Interaction
- 先根据文本描述生成物体形状先验,再通过视觉信息精细校准。
- 在Dex-YCB和Obman数据集上分别达到1.979和5.468的物体切比雪夫距离。
- 对遮挡鲁棒,兼容多种先验来源,适合真实复杂场景应用。
现有3D手物交互重建普遍依赖预定义的3D物体模板,但其获取需大量人工工作,且难以适应无约束交互场景(如严重遮挡)。为此,我们提出文本引导生成与精修框架TIGeR,利用直观的文本驱动先验来指导物体形状优化与位姿估计。该框架采用两阶段设计:第一阶段基于文本描述,使用现成模型生成形状先验,避免繁琐的3D建模;第二阶段通过2D-3D协同注意力机制,校正合成原型与真实物体之间的几何偏差。TIGeR在广泛使用的Dex-YCB和Obman数据集上分别取得1.979和5.468的物体切比雪夫距离,优于现有无模板方法。特别地,该框架对遮挡具有强鲁棒性,并可兼容异构先验源(如检索到的手工原型),适用于实际部署场景。
原文摘要 · Abstract (English)
Pre-defined 3D object templates are widely used in 3D reconstruction of hand-object interactions. However, they often require substantial manual efforts to capture or source, and inherently restrict the adaptability of models to unconstrained interaction scenarios, e.g., heavily-occluded objects. To overcome this bottleneck, we propose a new Text-Instructed Generation and Refinement (TIGeR) framework, harnessing the power of intuitive text-driven priors to steer the object shape refinement and pose estimation. We use a two-stage framework: a text-instructed prior generation and vision-guided refinement. As the name implies, we first leverage off-the-shelf models to generate shape priors according to the text description without tedious 3D crafting. Considering the geometric gap between the synthesized prototype and the real object interacted with the hand, we further calibrate the synthesized prototype via 2D-3D collaborative attention. TIGeR achieves competitive performance, i.e., 1.979 and 5.468 object Chamfer distance on the widely-used Dex-YCB and Obman datasets, respectively, surpassing existing template-free methods. Notably, the proposed framework shows robustness to occlusion, while maintaining compatibility with heterogeneous prior sources, e.g., retrieved hand-crafted prototypes, in practical deployment scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。