arXiv:2601.03667cs.CVcs.LG2026-01被引 1

用2D点轨迹提升手物交互识别,无需检测关键区域

TRec: Learning Hand-Object Interactions through 2D Point Track Motion

  • 随机采样图像点并跟踪其运动轨迹,作为额外时序线索
  • 仅需初始帧和点轨迹即显著提升识别准确率,优于无运动信息模型
  • 无需人体或物体检测,适合轻量级手物交互理解任务

我们提出一种新型手物交互动作识别方法,利用2D点轨迹作为额外的运动线索。现有方法多依赖RGB外观、人体姿态估计或二者结合,而本工作表明:对视频中随机采样的图像点进行跨帧跟踪,可显著提升识别性能。与以往方法不同,我们不检测手部、物体或交互区域,而是使用CoTracker在每段视频中追踪一组随机初始化的点,将得到的轨迹连同对应图像帧输入Transformer-based识别模型。令人惊讶的是,即使仅提供初始帧和点轨迹,不使用完整视频序列,该方法仍能实现明显性能提升。实验结果表明,引入2D点轨迹始终优于仅使用静态图像的相同模型,凸显其作为轻量但高效的手物交互理解表示的巨大潜力。

原文摘要 · Abstract (English)

We present a novel approach for hand-object action recognition that leverages 2D point tracks as an additional motion cue. While most existing methods rely on RGB appearance, human pose estimation, or their combination, our work demonstrates that tracking randomly sampled image points across video frames can substantially improve recognition accuracy. Unlike prior approaches, we do not detect hands, objects, or interaction regions. Instead, we employ CoTracker to follow a set of randomly initialized points through each video and use the resulting trajectories, together with the corresponding image frames, as input to a Transformer-based recognition model. Surprisingly, our method achieves notable gains even when only the initial frame and the point tracks are provided, without incorporating the full video sequence. Experimental results confirm that integrating 2D point tracks consistently enhances performance compared to the same model trained without motion information, highlighting their potential as a lightweight yet effective representation for hand-object action understanding.

动作识别点轨迹轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。