人类和模型无需训练即可识别去外观动作,关键靠运动信息。
Appearance-free Action Recognition: Zero-shot Generalization in Humans and a Two-Pathway Model

- 用双流3D CNN建模,融合颜色与运动流,引入共形门控机制。
- 在两种去外观视频上零样本识别准确率显著高于随机水平。
- 运动流对去外观泛化至关重要,适合研究视觉认知与视频模型。
动作识别是社交物种的基本能力,但其计算机制尚不明确。经典心理学实验显示,即使形状线索被破坏,人类仍能感知身体运动。近期研究使用真实动作视频及其去外观版本(保留运动但无静态形状)进行训练。本文通过实验室心理物理学实验测试22名参与者:先用自然动作视频(UCF5数据集)训练识别五类动作,再零样本测试其对两类去外观变换的识别能力——(i)现有数据集中的密集噪声运动视频(AFD5),(ii)随机点去外观视频。结果显示,参与者在两类去外观视频上的识别准确率均显著高于随机水平,尽管低于自然视频表现。为此,我们构建了一个基于3D CNN的双流模型,包含RGB(形态)流与光流(运动)流,并引入受格式塔共同命运分组启发的共形门控机制。该模型在两个去外观数据集上均实现良好零样本泛化,优于当前主流视频分类模型,接近人类表现。分析表明,运动流对去外观泛化起决定性作用,而形态流提升自然视频上的性能。研究强调运动表征在泛化中的重要性,支持多流架构用于视频动作识别建模。
原文摘要 · Abstract (English)
Action recognition is a fundamental ability for social species. Yet, its underlying computations are not well understood. Classical psychophysical studies using simplified stimuli have shown that humans can perceive body motion even under degradation of relevant shape cues. Recent work using real-world action videos and their appearance-free counterparts (that preserve motion but lack static shape cues) included explicit training of humans and models on the appearance-free videos. Whether humans and vision models generalize in a zero-shot manner to appearance-free transformations of real-world action videos is not yet known. To measure this generalization in humans, we conducted a laboratory-based psychophysics experiment. 22 participants were trained to recognize five action categories using naturalistic videos (UCF5 dataset), and tested zero-shot on two types of appearance-free transformations: (i) dense-noise motion videos from an existing dataset (AFD5) and (ii) random-dot appearance-free videos. We find that participants recognize actions in both types of appearance-free videos well above chance, albeit with reduced accuracy compared to naturalistic videos. To model this behavior, we developed a two-pathway 3D CNN-based model combining an RGB (form) stream and an optical flow (motion) stream, including a coherence-gating mechanism inspired by Gestalt common-fate grouping. Our model generalizes to both appearance-free datasets and outperforms contemporary video classification models, narrowing the gap to human performance. We find that the motion pathway is critical for generalization to appearance-free videos, while the form pathway improves performance on naturalistic videos. Our findings highlight the importance of motion-based representations for generalization to appearance-free videos, and support the use of multi-stream architectures to model video-based action recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。