实时保护3D视频隐私,防止敏感内容在多视角融合中泄露。
Cloak of Invisibility: Real-Time Privacy-Preserving Volumetric Video Streaming

- 多视角同步检测并掩蔽私密物体,结合深度感知避免几何信息泄露。
- 真实场景下掩蔽精度达Dice=0.792,Recall=0.908,保持30 FPS以上实时流传输。
- 适用于远程会议、沉浸式教育等需保护隐私的3D视觉应用。
体积视频流将隐私问题转化为三维多视角难题。与普通视频可逐帧擦除不同,RGB-D体积管道从多摄像头捕获人物、房间和私人物品,并融合为共享3D表示。若某视角遗漏或部分删除私密对象,其仍可能在重建场景中重现。这给3D远程呈现、教育、娱乐及沉浸式应用带来隐私挑战:敏感内容应在摄像头端移除,同时保留公开部分以支持实时重建。现有体积流系统主要优化重建质量、数据传输与延迟,而隐私保护方法针对单摄像头单帧图像,无法应对校准后的多视角RGB-D融合。本文提出InViStream,一种面向该场景的实时“源端隐私保护”系统。InViStream解决三大挑战:私密物体在各视角表现不同、仅颜色掩蔽会遗留深度隐私泄漏、同类公共/私密实例需在云端融合前一致区分。为此,InViStream结合目标检测与深度感知掩蔽,跨校准视角传播公/私决策,仅融合清洗后的点云。在合成与真实RGB-D场景(办公室、会议室、客厅及多人多物场景)上评估,InViStream在合成数据上获得Dice=0.799、Recall=0.891,真实数据上为Dice=0.792、Recall=0.908,合成SSIM>0.98,实现实时流传输超过30 FPS。
原文摘要 · Abstract (English)
Volumetric video streaming turns privacy into a 3D, multi-view problem. Unlike ordinary video, where sensitive content can often be redacted frame by frame, RGB-D volumetric pipelines capture people, rooms, and personal objects from multiple cameras and fuse them into a shared 3D representation. A private object missed in one view, or only partially removed before fusion, can therefore reappear in the reconstructed scene. This creates a privacy challenge for 3D telepresence, education, entertainment, and immersive applications: private content should be removed before raw visual and geometric data leave the camera side, while the public part of the scene should remain useful for real-time reconstruction. Existing volumetric streaming systems mainly optimize reconstruction, data movement, and latency, while privacy-preserving vision methods are designed for single-camera, single-frame images and do not directly address calibrated multi-view RGB-D fusion. We present InViStream, a real-time "privacy-from-source" system designed for this setting. InViStream addresses three challenges in volumetric capture: private objects may appear differently across views, RGB masking alone can leave geometric privacy leakage in depth, and public/private instances of the same class must be separated consistently before cloud-side fusion. To address these challenges, InViStream combines object detection with depth-aware masking, propagates public/private decisions across calibrated views, and fuses only sanitized point clouds. We evaluate InViStream on synthetic and real RGB-D scenes, including offices, conference rooms, living rooms, and settings with multiple public and private people and objects. InViStream achieves synthetic Dice/Recall of 0.799/0.891 and real Dice/Recall of 0.792/0.908, with synthetic SSIM above 0.98 and real-time streaming above 30 FPS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。