arXiv:2508.00085cs.CVcs.AI2025-08ICCV被引 1

测试动作识别模型在新情境下的迁移能力,发现性能大幅下降。

Punching Bag vs. Punching Person: Motion Transferability in Videos

  • 构建三个数据集,评估模型跨场景动作识别能力。
  • 模型在新动作上准确率下降超30%,尤其在细粒度任务中。
  • 大模型更依赖空间线索,但难应对复杂时间推理。

动作识别模型虽具强泛化能力,但在不同情境下能否有效迁移高层运动概念仍不明确。例如,模型能否识别未见过的“打人”动作?为此,本文引入运动可迁移性框架,构建三个数据集:(1) Syn-TA,含3D物体运动的合成数据;(2) Kinetics400-TA;(3) Something-Something-v2-TA,均源自自然视频数据集并进行适配。评估13个先进模型后发现,模型在新情境中识别高层动作时性能显著下降。分析表明:(1) 多模态模型对细粒度未知动作的适应性差于粗粒度动作;(2) 无偏见的Syn-TA与真实世界数据集同样具有挑战性,控制环境下模型性能下降更明显;(3) 更大模型在空间线索主导时提升迁移能力,但在需密集时间推理时表现不佳,过度依赖物体与背景线索会阻碍泛化。进一步探索解耦粗粒度与细粒度运动对时序挑战数据集的提升效果。本研究为评估动作识别中的运动可迁移性建立了关键基准。数据集与代码已开源。

原文摘要 · Abstract (English)

Action recognition models demonstrate strong generalization, but can they effectively transfer high-level motion concepts across diverse contexts, even within similar distributions? For example, can a model recognize the broad action "punching" when presented with an unseen variation such as "punching person"? To explore this, we introduce a motion transferability framework with three datasets: (1) Syn-TA, a synthetic dataset with 3D object motions; (2) Kinetics400-TA; and (3) Something-Something-v2-TA, both adapted from natural video datasets. We evaluate 13 state-of-the-art models on these benchmarks and observe a significant drop in performance when recognizing high-level actions in novel contexts. Our analysis reveals: 1) Multimodal models struggle more with fine-grained unknown actions than with coarse ones; 2) The bias-free Syn-TA proves as challenging as real-world datasets, with models showing greater performance drops in controlled settings; 3) Larger models improve transferability when spatial cues dominate but struggle with intensive temporal reasoning, while reliance on object and background cues hinders generalization. We further explore how disentangling coarse and fine motions can improve recognition in temporally challenging datasets. We believe this study establishes a crucial benchmark for assessing motion transferability in action recognition. Datasets and relevant code: https://github.com/raiyaan-abdullah/Motion-Transfer.

动作识别迁移能力视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。