用深度学习自动选关键帧,让零售视频标注效率翻倍、成本减半。
Efficient Retail Video Annotation: A Robust Key Frame Generation Approach for Product and Customer Interaction Analysis
- 用神经网络提取视频特征,结合商品检测技术自动选关键帧。
- 标注准确率接近人工水平,节省约50%人力成本。
- 适合需要高效分析顾客行为的零售企业使用。
精准的视频标注在现代零售应用中至关重要,涵盖顾客行为分析、商品互动检测和店内活动识别。然而,传统方法严重依赖耗时的人工标注,导致帧选择不稳健且运营成本高。为解决零售领域的挑战,我们提出一种基于深度学习的方法,自动识别零售视频中的关键帧,并实现商品与顾客的自动标注。该方法利用深度神经网络,通过嵌入视频帧并结合针对零售环境优化的目标检测技术,学习判别性特征。实验结果表明,该方法优于传统方法,在标注准确率上接近人工标注水平的同时,显著提升了整体效率。值得注意的是,该方法可使视频标注成本平均降低约50%。通过仅需人工验证或调整不足5%的检测帧,其余帧可全自动标注且不降低质量,零售商能大幅降低运营成本。关键帧自动检测极大节省了标注时间与人力,对购物旅程分析、商品互动检测和店内安全监控等应用具有极高价值。
原文摘要 · Abstract (English)
Accurate video annotation plays a vital role in modern retail applications, including customer behavior analysis, product interaction detection, and in-store activity recognition. However, conventional annotation methods heavily rely on time-consuming manual labeling by human annotators, introducing non-robust frame selection and increasing operational costs. To address these challenges in the retail domain, we propose a deep learning-based approach that automates key-frame identification in retail videos and provides automatic annotations of products and customers. Our method leverages deep neural networks to learn discriminative features by embedding video frames and incorporating object detection-based techniques tailored for retail environments. Experimental results showcase the superiority of our approach over traditional methods, achieving accuracy comparable to human annotator labeling while enhancing the overall efficiency of retail video annotation. Remarkably, our approach leads to an average of 2 times cost savings in video annotation. By allowing human annotators to verify/adjust less than 5% of detected frames in the video dataset, while automating the annotation process for the remaining frames without reducing annotation quality, retailers can significantly reduce operational costs. The automation of key-frame detection enables substantial time and effort savings in retail video labeling tasks, proving highly valuable for diverse retail applications such as shopper journey analysis, product interaction detection, and in-store security monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。