arXiv:2501.07100cs.CVcs.AI2025-01AAAI被引 2

用超二次曲面表示物体,提升手物交互与组合动作识别的泛化能力。

Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics

论文配图:Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics
图 1 · 摘自论文原文
  • 用超二次曲面替代3D框,实现无需模板的物体重建
  • 在组合动作识别上超越现有方法,尤其在未见物体上表现更优
  • 适合关注3D交互理解与模型泛化的研究者

随着第一人称视角3D手物交互数据集的出现,统一建模手物姿态估计与动作识别成为研究热点。然而,现有方法在未见物体上识别已见动作时仍存在局限,主要因3D边界框难以有效表征物体形状与运动。同时,测试时依赖物体模板也限制了模型泛化能力。为此,本文提出采用超二次曲面作为替代性3D物体表示,并在无模板物体重建与动作识别任务中验证其有效性。此外,我们发现纯视觉外观方法已能超越统一模型,表明3D几何信息的优势尚不明确。因此,我们引入更具挑战性的组合性动作识别任务,即训练与测试集中的动词-名词组合无重叠。我们扩展了H2O与FPHA数据集并构建新型协同学习框架,可显式建模手与物体间的几何关系。通过大量定量与定性评估,结果表明该方法在(组合)动作识别上显著优于现有最优模型。

原文摘要 · Abstract (English)

With the availability of egocentric 3D hand-object interaction datasets, there is increasing interest in developing unified models for hand-object pose estimation and action recognition. However, existing methods still struggle to recognise seen actions on unseen objects due to the limitations in representing object shape and movement using 3D bounding boxes. Additionally, the reliance on object templates at test time limits their generalisability to unseen objects. To address these challenges, we propose to leverage superquadrics as an alternative 3D object representation to bounding boxes and demonstrate their effectiveness on both template-free object reconstruction and action recognition tasks. Moreover, as we find that pure appearance-based methods can outperform the unified methods, the potential benefits from 3D geometric information remain unclear. Therefore, we study the compositionality of actions by considering a more challenging task where the training combinations of verbs and nouns do not overlap with the testing split. We extend H2O and FPHA datasets with compositional splits and design a novel collaborative learning framework that can explicitly reason about the geometric relations between hands and the manipulated object. Through extensive quantitative and qualitative evaluations, we demonstrate significant improvements over the state-of-the-arts in (compositional) action recognition.

3D重建动作识别超二次曲面协同学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。