arXiv:2606.07171cs.CV2026-06

提出首个关注编辑后恢复的隐私保护框架,解决云编辑中替换内容无法还原的问题。

When Recovery Matters: The Blind Spot of Surrogate Privacy in MLLM Editing

论文配图:When Recovery Matters: The Blind Spot of Surrogate Privacy in MLLM Editing
图 1 · 摘自论文原文
  • 设计双任务评估体系:判断替换内容能否实现原图编辑效果,及能否恢复原始图像。
  • 提出ERMA与C2E-S2SER方法,在36类隐私数据上提升13.9%编辑可实现性预测准确率。
  • 适合研究多模态模型隐私保护、图像编辑恢复的学者,尤其关注真实场景还原能力。

多模态大语言模型支持灵活的指令驱动图像编辑,但用户图像可能暴露多样且个人化的隐私内容。主流隐私保护策略通常在云端编辑前用替代内容替换敏感区域,导致输出为编辑后的替代图像而非期望的原始图像,忽视了本地恢复的设计与评估。为此,我们提出SPPE(基于替代内容的隐私保护编辑)基准,涵盖36个细粒度隐私类别和65种编辑指令,定义两个互补任务:1)可编辑性评估,即在云端交互前估计替代内容是否能产生与原始图像一致的编辑结果;2)替代到源图像的编辑恢复,评估编辑后的替代内容能否还原至原始私有图像并保留编辑效果。针对每个任务,我们分别提出相应方法:ERMA通过指令感知的多模态关系建模预测替代内容可编辑性; method则利用替代编辑对作为视觉编辑证据,以源图像为保持原始结构的锚点,实现循环一致性恢复。在SPPE与InstructPix2Pix上的实验显示,两项任务均取得一致改进。对于可编辑性评估,ERMA相比最优基线在SRCC上提升13.9%,在PLCC上提升12.3%。对于替代到源图像的编辑恢复,C2E-S2SER在所有8项源完整性与编辑一致性指标上均优于SOER。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) enable flexible instruction-driven image editing, but privacy risks arise when user images expose diverse and user-specific private content. Canonical privacy protection strategies typically substitute sensitive regions with surrogate content before cloud editing. Yet, the resulting output is often an edited surrogate rather than the desired edited source image, neglecting the local recovery in both design and evaluation scope. To this end, we introduce SPPE (Surrogate-based Privacy-Preserving Editing), the first recovery-oriented benchmark covering 36 fine-grained privacy categories and 65 editing instructions. It defines two complementary tasks: 1) editability assessment, which estimates before cloud interaction whether a surrogate can induce an edit consistent with the original image; and 2) surrogate-to-source edit recovery, which evaluates whether the edited surrogate can be transferred back to the private source with the edit effect preserved. We address each task with a dedicated method: ERMA predicts surrogate editability through instruction-aware multimodal relation modeling, while \method performs cycle-consistent recovery by using the surrogate editing pair as visual edit evidence and the source image as a source-preserving anchor. Experiments on SPPE and InstructPix2Pix show consistent improvements on both tasks. For editability assessment, ERMA improves over the best-performing baselines by 13.9% in SRCC and 12.3% in PLCC. For surrogate-to-source edit recovery, C2E-S2SER outperforms SOER across all 8 source integrity and edit consistency metrics on SPPE.

隐私保护图像编辑多模态恢复机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。