arXiv:2603.29631cs.CVcs.DC2026-03

用新颖性过滤减少冗余帧,提升边缘摄像头跨模态检索效果

Storing Less, Finding More: How Novelty Filtering Improves Cross-Modal Retrieval on Edge Cameras

  • 在设备端用epsilon-net筛选语义新帧,构建干净嵌入索引
  • 仅用800万参数模型即达45.6% Hit@5,功耗低至2.7mW
  • 适合资源受限的实时视频检索场景,尤其适用边缘计算

持续运行的边缘摄像头产生连续视频流,冗余帧会因挤占结果导致跨模态检索性能下降。本文提出一种流式检索架构:设备端的epsilon-net过滤器仅保留语义新颖帧,构建去噪嵌入索引;跨模态适配器与云端重排序模块弥补紧凑编码器对齐能力不足。单次遍历的流式过滤器在两个第一人称数据集(AEA、EPIC-KITCHENS)上优于离线方法(k-means、远点采样、均匀采样、随机采样),涵盖800万至6.32亿参数的八种视觉-语言模型。整体架构在使用800万参数设备端编码器时,于留出数据上达到45.6% Hit@5,预计功耗仅2.7mW。

原文摘要 · Abstract (English)

Always-on edge cameras generate continuous video streams where redundant frames degrade cross-modal retrieval by crowding correct results out of top-k search. This paper presents a streaming retrieval architecture: an on-device epsilon-net filter retains only semantically novel frames, building a denoised embedding index; a cross-modal adapter and cloud re-ranker compensate for the compact encoder's weak alignment. A single-pass streaming filter outperforms offline alternatives (k-means, farthest-point, uniform, random) across eight vision-language models (8M-632M) on two egocentric datasets (AEA, EPIC-KITCHENS). Combined, the architecture reaches 45.6% Hit@5 on held-out data using an 8M on-device encoder at an estimated 2.7 mW.

边缘计算跨模态检索视频压缩节能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。