arXiv:2606.15335cs.CLcs.AI2026-06

用解耦表征保护跨组织文本协作中的隐私,防泄露更彻底。

Privacy-Preserving Text Sanitization for Distributed Agents Collaboration via Disentangled Representations

论文配图:Privacy-Preserving Text Sanitization for Distributed Agents Collaboration via Disentangled Representations
图 1 · 摘自论文原文
  • 将文本分解为任务语义与身份风格两部分,分别处理
  • 隐私泄露降低20倍,答案准确率仍保持83%
  • 适合需要跨机构协作又怕数据外泄的场景

当分布式代理在组织间交换文本时,隐私泄露不仅来自显式标识符,还源于格式习惯、词汇选择和句法模式等分布特征。我们提出DiSan(解耦净化)框架,作为Intern-Shannon多智能体协作系统的一部分。DiSan采用双流编码器,将文本分解为保留任务语义的源无关角色子空间和仅本地存在的源识别风格子空间。通过联邦原型对齐与对抗正则化实现联合训练,无需集中原始文本。实验表明:仅掩码19.2%的词元,TF-IDF风格归因仅降低18.6%;而DiSan使答案级敏感信息暴露减少20倍,在分布式多智能体RAG基准上保持83%答案忠实度,且在Enron数据集上使TF-IDF风格归因下降73.2%,神经探测器归因下降70.6%。

原文摘要 · Abstract (English)

When distributed agents exchange text across organizational boundaries, privacy leakage arises not only from explicit identifiers but also from distributional signatures such as formatting conventions, vocabulary choices, and syntactic patterns. We propose DiSan(Disentangled Sanitization), a privacy-preserving sanitization framework and a built-in component of Intern-Shannon for multi-agent collaboration. DiSan uses a two-stream encoder to factorize text into a source-invariant role subspace that preserves task semantics and a source-identifying style subspace that remains local. Federated proto-type alignment and adversarial regularization enable joint training without centralizing raw text. Experiments show that identifier-level masking is insufficient: masking 19.2% of tokens reduces TF-IDF stylometric attribution by only 18.6%. By contrast, DiSan reduces answer-level PII exposure by 20 times while maintaining 83% answer faithfulness on a distributed multi-agent RAG benchmark, and lowers Enron stylometric attribution by 73.2% under TF-IDF and 70.6% under a neural probe.

隐私保护文本净化多智能体解耦表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。