arXiv:2506.03171cs.CVcs.AI2025-06被引 2

在手机等设备上实时生成个性化视频摘要,保护隐私且无需上传数据。

EdgeVidSum: Real-Time Personalized Video Summarization at the Edge

  • 用缩略图容器替代逐帧处理,大幅降低计算量。
  • 轻量级2D CNN模型从缩略图中识别用户偏好内容并生成快进时间戳。
  • 可在Jetson Nano等资源受限设备运行,适合个人视频快速浏览。

EdgeVidSum是一种轻量级方法,可直接在边缘设备上生成长视频的个性化快进摘要。该方法通过创新的缩略图技术与高效神经架构实现实时视频摘要,同时保障用户隐私——所有数据本地处理。不同于传统逐帧分析方式,该方法使用缩略图容器显著降低计算复杂度,同时保持语义相关性。系统采用分层分析策略:轻量级2D CNN模型从缩略图中识别用户偏好的内容,并生成时间戳以构建快进摘要。交互式演示表明,该系统能为电影、体育赛事和电视剧等长视频生成个性化摘要。整个计算过程在资源受限设备(如Jetson Nano)上无缝完成,有效应对现代视频消费场景中的计算效率、个性化与隐私保护挑战。

原文摘要 · Abstract (English)

EdgeVidSum is a lightweight method that generates personalized, fast-forward summaries of long-form videos directly on edge devices. The proposed approach enables real-time video summarization while safeguarding user privacy through local data processing using innovative thumbnail-based techniques and efficient neural architectures. Unlike conventional methods that process entire videos frame by frame, the proposed method uses thumbnail containers to significantly reduce computational complexity without sacrificing semantic relevance. The framework employs a hierarchical analysis approach, where a lightweight 2D CNN model identifies user-preferred content from thumbnails and generates timestamps to create fast-forward summaries. Our interactive demo highlights the system's ability to create tailored video summaries for long-form videos, such as movies, sports events, and TV shows, based on individual user preferences. The entire computation occurs seamlessly on resource-constrained devices like Jetson Nano, demonstrating how EdgeVidSum addresses the critical challenges of computational efficiency, personalization, and privacy in modern video consumption environments.

视频摘要边缘计算个性化隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。