arXiv:2603.27811cs.CVcs.LG2026-03被引 1

通过加密流量推断位置,无需访问原始视频数据。

Tracking without Seeing: Geospatial Inference using Encrypted Traffic from Distributed Nodes

  • 从加密包大小推断画面变化,间接感知物体运动。
  • 在无原始信号条件下实现2.33米追踪误差,接近物体尺寸。
  • 适合隐私敏感场景下的远程监控与协同感知系统。

传统动态环境观测依赖多源分布式传感器的原始信号融合。本文提出一种新方法:仅利用加密的报文级信息完成地理空间推断,无需访问原始传感数据。我们进一步研究如何将这种间接信息与可获得的直接传感数据融合以扩展推断能力。提出GraySense框架,通过分析无法访问的摄像头无线视频流中加密报文的包大小等特征,实现对移动目标的地理空间跟踪。该框架包含两阶段:(1) 报文分组模块,从加密网络流量中识别帧边界并估计帧大小;(2) 基于带循环状态的Transformer编码器的跟踪器,融合基于报文的间接输入和可选的直接摄像头输入以估计目标位置。在CARLA模拟器生成的真实视频及不同网络条件下的仿真环境中进行大量实验,结果表明,在无原始信号访问的情况下,灰度感知框架达到2.33米(欧氏距离)的追踪误差,小于被追踪对象尺寸(4.61米×1.93米)。据我们所知,这是首次实现此类能力,拓展了隐式信号在传感中的应用边界。

原文摘要 · Abstract (English)

Accurate observation of dynamic environments traditionally relies on synthesizing raw, signal-level information from multiple distributed sensors. This work investigates an alternative approach: performing geospatial inference using only encrypted packet-level information, without access to the raw sensory data. We further explore how this indirect information can be fused with directly available sensory data to extend overall inference capabilities. We introduce GraySense, a learning-based framework that performs geospatial object tracking by analyzing encrypted wireless video transmission traffic, such as packet sizes, from cameras with inaccessible streams. GraySense leverages the inherent relationship between scene dynamics and transmitted packet sizes to infer object motion. The framework consists of two stages: (1) a Packet Grouping module that identifies frame boundaries and estimates frame sizes from encrypted network traffic, and (2) a Tracker module, based on a Transformer encoder with a recurrent state, which fuses indirect packet-based inputs with optional direct camera-based inputs to estimate the object's position. Extensive experiments with realistic videos from the CARLA simulator and emulated networks under varying conditions show that GraySense achieves 2.33 meters tracking error (Euclidean distance) without raw signal access, within the dimensions of tracked objects (4.61m x 1.93m). To our knowledge, this capability has not been previously demonstrated, expanding the use of latent signals for sensing.

位置推断加密流量视频跟踪联邦感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。