用结构信息重生成人体,实现真实、连贯且匿名的全身视频处理。
ReGenHuman: Re-Generating Human Appearances for Realistic Full-Body Video Anonymization

- 不修改原图,而是基于姿态、分割和深度信息重建人体
- 在隐私、画质和下游任务效果上均优于现有方法
- 适合需要保护隐私又保留视频可用性的研究与应用
人体中心视频的匿名化问题尚未得到充分研究。以往方法要么模糊或遮盖像素,牺牲真实感和后续用途;要么逐帧生成,导致时间不连贯。我们提出 ReGenHuman,首个同时具备真实性、时间一致性与匿名性的全身视频匿名化流程。不同于直接编辑输入的方法,我们采用‘重生成而非编辑’范式,将2D姿态、分割和单目深度组合为两个互补的条件流——StructAll 和 StructHuman,用于微调视频到视频的扩散模型,完全基于无身份的结构线索合成人体区域。我们在隐私性、画质和实用性方面评估模型,结果表明 ReGenHuman 在三项指标间取得了最优平衡。此外,我们的匿名视频仍可用于下游任务,如视频问答。
原文摘要 · Abstract (English)
Anonymizing human-centric video data is an understudied problem. Prior anonymization techniques either blur or redact pixels at the cost of realism and downstream utility, or generate frame-by-frame at the cost of temporal coherence. We introduce ReGenHuman, the first full-body video anonymization pipeline that is simultaneously realistic, temporally consistent, and anonymous by construction. Contrary to past approaches which redact or edit the inputs directly, we propose a regenerate, don't edit paradigm. Our approach composites 2D pose, segmentation, and monocular depth into two complementary conditioning streams - StructAll and StructHuman, which are used to fine-tune a video-to-video diffusion backbone on in-the-wild human videos, synthesizing the human regions entirely from identity-free structural cues. We evaluate our model on privacy, quality, and utility, and show that our ReGenHuman achieves the best tradeoff across all three axes against current baselines. We further show that our anonymized videos remain effective for downstream tasks, including video question answering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。