仅用热成像实现隐私保护下的人群计数,告别持续拍摄可见光视频
Thermal-Only Crowd Counting with Deployment-Time Privacy Protection

- 用深度到可见光的扩散模型做跨模态桥梁,增强热成像特征表达
- 单步LCM去噪保留结构信息,多步会引入误差导致计数不准
- 推理时只需热成像,实测性能媲美多模态融合方法
尽管RGB-热成像人群计数展现出潜力,但该范式面临两大瓶颈:可见光数据引发公共监控中的隐私担忧,多模态错位降低融合性能。本文提出首个专为隐私敏感场景设计的纯热成像人群计数框架,在推理阶段完全消除对可见光数据的依赖,显著降低公共部署中持续采集可见光图像带来的隐私风险。为缓解热成像的语义模糊性,我们利用深度到可见光的扩散模型作为跨模态桥梁,提取更具判别性的特征以增强热成像表征。关键发现:单步LCM去噪生成的特征最忠实于深度条件信号的结构内容,而多步方法会逐步脱离条件输入并累积误差,损害计数精度。在RGBT-CC和DroneRGBT数据集上的实验表明,本方法在仅使用热成像输入的情况下,性能可与最先进的RGB-T融合方法相媲美,同时避免了实际部署中持续采集可见光图像的核心隐私问题。代码将公开。
原文摘要 · Abstract (English)
While RGB-Thermal crowd counting has shown promise, the paradigm faces critical limitations: RGB data raises privacy concerns in public surveillance, and multi-modal misalignment degrades fusion performance. We propose the first thermal-only framework specifically designed for privacy-conscious crowd counting, eliminating RGB dependency at inference time and substantially reducing the privacy exposure associated with continuous RGB capture in public surveillance deployments. To mitigate thermal ambiguity, we leverage depth-to-RGB diffusion models as a cross-modal bridge, extracting discriminative features that enhance thermal representations. Critically, we demonstrate that single-step LCM denoising yields features most faithful to the structural content of the depth conditioning signal, while multi-step approaches progressively decouple features from the conditioning input and accumulate errors that degrade counting accuracy. Experiments on RGBT-CC and DroneRGBT datasets show our method achieves competitive performance against state-of-the-art RGB-T fusion methods, while requiring only thermal input during inference, eliminating the need for continuous RGB capture that constitutes the primary privacy concern in real-world surveillance deployment. The code will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。