arXiv:2502.03459cs.CV2025-02AAAI被引 6

将3D骨骼信息融入视觉语言模型,提升日常活动理解能力

SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living

  • 用骨骼-语言模型融合3D骨骼与文本信息,构建联合嵌入空间
  • 在3个主流ADL数据集上实现零样本动作识别与视频描述生成
  • 推理时无需骨骼数据,适合真实场景部署

CLIP等视觉语言模型推动了可泛化至未见视频和动作的基础视频模型发展。然而,这些模型通常基于网络视频训练,难以捕捉日常生活活动(ADL)视频中的挑战,如外观相似、运动微弱和多视角问题。现有方法通过结合3D骨骼与RGB视频应对这些挑战,但缺乏语言整合,限制了对未见动作类别的泛化能力。本文提出SKI模型,将3D骨骼信息融入视觉语言嵌入空间。SKI模型通过协同训练引入骨骼-语言模型SkeletonCLIP,将骨骼信息注入视觉语言模型(VLM)和大视觉语言模型(LVLM)。值得注意的是,SKI模型在推理阶段无需骨骼数据,增强了实际应用的鲁棒性。其有效性在三个主流ADL数据集上通过零样本动作识别与视频字幕生成任务得到验证。

原文摘要 · Abstract (English)

The introduction of vision-language models like CLIP has enabled the development of foundational video models capable of generalizing to unseen videos and human actions. However, these models are typically trained on web videos, which often fail to capture the challenges present in Activities of Daily Living (ADL) videos. Existing works address ADL-specific challenges, such as similar appearances, subtle motion patterns, and multiple viewpoints, by combining 3D skeletons and RGB videos. However, these approaches are not integrated with language, limiting their ability to generalize to unseen action classes. In this paper, we introduce SKI models, which integrate 3D skeletons into the vision-language embedding space. SKI models leverage a skeleton-language model, SkeletonCLIP, to infuse skeleton information into Vision Language Models (VLMs) and Large Vision Language Models (LVLMs) through collaborative training. Notably, SKI models do not require skeleton data during inference, enhancing their robustness for real-world applications. The effectiveness of SKI models is validated on three popular ADL datasets for zero-shot action recognition and video caption generation tasks.

视觉语言模型日常活动理解3D骨骼零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。