让AI视频分析更快更准,边云协同处理用户实时提示
SAMEdge: An Edge-cloud Video Analytics Architecture for the Segment Anything Model
- 边云协同架构,动态分配计算任务
- 视觉提示转换算法降低传输开销,提升响应速度
- 适合需要实时交互的智能视频应用开发者
随着人工智能发展,单一大型模型已能处理多样视频分析任务。其中关键基础技术是分割一切模型(SAM),可根据用户输入的提示动态调整分析内容。然而,在边缘设备资源受限的情况下,实现实时响应对用户体验至关重要,尤其当用户持续添加或修改提示时。本文提出SAMEdge,一种面向边缘用户的边云协同视频分析架构,用于支持SAM计算。该架构在边缘与云端集成新模块,以在延迟约束下最大化基于视觉提示和图像提示输入的分析准确率。通过引入视觉提示转换算法缓解提示编码资源压力,并采用高效工作负载划分解决图像编码挑战。SAMEdge基于Meta AI开源的SAM项目实现。通过一个视觉导览应用案例展示其实际价值。评估结果表明,在不同网络带宽下,SAMEdge显著提升了视频分析应用的准确性。
原文摘要 · Abstract (English)
As artificial intelligence continues to evolve, it is increasingly capable of handling a wide range of video analytics tasks with merely one large model. One of the key foundation technologies is the Segment Anything Model (SAM), which allows the video analytics tasks to be determined on the fly according to the input prompts from the user. However, achieving real-time response in video analytics applications is crucial for user experiences due to the limited communication and computation resources on the edge, especially with SAM, where users may continuously interact by adding or adjusting prompts. In this paper, we propose SAMEdge, a novel edge-cloud computing architecture designed to support SAM computations for edge users. SAMEdge integrates new modules on the edge and the cloud to maximize analytics accuracy under visual prompts and image prompts input with latency constraints. It addresses resource challenges associated with prompt encoding and image encoding by offering a visual prompt transformation algorithm for visual prompts and efficient workload partitioning for image encoding. SAMEdge is implemented by extending the open-source SAM project from Meta AI. We demonstrate the practical application of SAMEdge through a case study on a Visual Tour Guide application. Our evaluation indicates that SAMEdge significantly enhances the accuracy of the video analytics application under distinct network bandwidths across various prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。