arXiv:2506.13458cs.CVcs.CL2025-06

用图文预训练提升单图动作识别准确率

Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images

  • 利用图文对比学习预训练模型
  • 单图动作识别准确率达76%
  • 适合智能监控与辅助系统应用

在单张图像中识别人类活动可支持索引、安全与辅助应用,但缺乏运动线索。基于285张标注为行走、跑步、坐、站的MSCOCO图像,原始CNN模型准确率为41%。微调多模态CLIP后,准确率提升至76%,表明对比视觉-语言预训练显著提升真实场景下单图动作识别性能。

原文摘要 · Abstract (English)

Recognising human activity in a single photo enables indexing, safety and assistive applications, yet lacks motion cues. Using 285 MSCOCO images labelled as walking, running, sitting, and standing, scratch CNNs scored 41% accuracy. Fine-tuning multimodal CLIP raised this to 76%, demonstrating that contrastive vision-language pre-training decisively improves still-image action recognition in real-world deployments.

动作识别图文预训练单图理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。