arXiv:2506.13458cs.CVcs.CL2025-06
用图文预训练提升单图动作识别准确率
Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images
- 利用图文对比学习预训练模型
- 单图动作识别准确率达76%
- 适合智能监控与辅助系统应用
在单张图像中识别人类活动可支持索引、安全与辅助应用,但缺乏运动线索。基于285张标注为行走、跑步、坐、站的MSCOCO图像,原始CNN模型准确率为41%。微调多模态CLIP后,准确率提升至76%,表明对比视觉-语言预训练显著提升真实场景下单图动作识别性能。
原文摘要 · Abstract (English)
Recognising human activity in a single photo enables indexing, safety and assistive applications, yet lacks motion cues. Using 285 MSCOCO images labelled as walking, running, sitting, and standing, scratch CNNs scored 41% accuracy. Fine-tuning multimodal CLIP raised this to 76%, demonstrating that contrastive vision-language pre-training decisively improves still-image action recognition in real-world deployments.
动作识别图文预训练单图理解
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。