arXiv:2601.19850cs.CV2026-01中稿 · ICLR被引 4

用上下文学习提升第一视角手部三维重建的准确性与泛化能力

EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning

  • 基于视觉语言模型检索互补样本,增强上下文理解
  • 在ARCTIC和EgoExo4D上超越现有方法,提升手物交互推理能力
  • 适合需要高精度手部重建的智能交互与虚拟现实应用

第一视角视觉中的鲁棒3D手部重建面临深度模糊、自遮挡和复杂手物交互等挑战。以往方法通过扩充训练数据或引入辅助线索缓解问题,但在未见场景中表现不佳。本文提出EgoHandICL,首个用于3D手部重建的上下文学习框架,通过视觉语言模型引导的互补样本检索、专为多模态上下文设计的分词器,以及基于掩码自编码器的架构,结合手部引导的几何与感知目标进行训练。在ARCTIC和EgoExo4D数据集上的实验表明,该方法持续优于当前最优模型。此外,我们验证了其在真实场景中的泛化能力,并通过重建的手部作为视觉提示,提升了EgoVLM对手物交互的推理性能。代码与数据已开源。

原文摘要 · Abstract (English)

Robust 3D hand reconstruction in egocentric vision is challenging due to depth ambiguity, self-occlusion, and complex hand-object interactions. Prior methods mitigate these issues by scaling training data or adding auxiliary cues, but they often struggle in unseen contexts. We present EgoHandICL, the first in-context learning (ICL) framework for 3D hand reconstruction that improves semantic alignment, visual consistency, and robustness under challenging egocentric conditions. EgoHandICL introduces complementary exemplar retrieval guided by vision-language models (VLMs), an ICL-tailored tokenizer for multimodal context, and a masked autoencoder (MAE)-based architecture trained with hand-guided geometric and perceptual objectives. Experiments on ARCTIC and EgoExo4D show consistent gains over state-of-the-art methods. We also demonstrate real-world generalization and improve EgoVLM hand-object interaction reasoning by using reconstructed hands as visual prompts. Code and data: https://github.com/Nicous20/EgoHandICL

3D手部重建上下文学习第一视角手物交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。