arXiv:2605.20258cs.LGcs.AI2026-05

让大模型在不牺牲任务能力的前提下,更智能地决定何时该说、何时不该说。

It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs

论文配图:It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs
图 1 · 摘自论文原文
  • 用双教师自蒸馏分离信息隐藏与任务处理,实现隐私与性能解耦
  • 在不依赖外部监督的情况下,优于强化学习等主流方法的披露控制效果
  • 特别适合需要长期记忆敏感信息的智能体工作流场景

上下文完整性(CI)定义隐私不仅在于隐藏信息,更在于依据特定语境规范信息流动。随着大语言模型作为个人代理处理敏感任务,遵循CI至关重要。然而,即使前沿模型在披露决策上仍不可靠,现有缓解策略常以降低任务性能为代价。为此,我们提出SELFCI——一种互补式自蒸馏框架,将信息抑制与任务求解解耦。SELFCI通过联合优化两个独立教师分布的反向KL散度:一个保留任务相关性以保障实用性,另一个强制最小且恰当的披露。该互补设计生成乘积专家(PoE)目标,使策略对齐能力与隐私要求的交集。实证表明,无须昂贵外部监督,SELFCI持续优于如GRPO等先进基线方法。这一优势在涉及智能体工作流和累积私密上下文的跨域场景中依然成立,表明SELFCI为实现上下文完整性对齐提供了可行路径。

原文摘要 · Abstract (English)

Contextual Integrity (CI) defines privacy not merely as keeping information hidden, but as governing information flows according to the norms of a given context. As large language models are increasingly deployed as personal agents handling sensitive workflows, adhering to CI becomes critical. However, even frontier models remain unreliable in making disclosure decisions, and existing mitigation strategies often degrade underlying task performance. To overcome this privacy-utility trade-off, we propose SELFCI, a complementary self-distillation framework that decouples information suppression from task resolution. SELFCI jointly optimizes two independent reverse KL divergences over distinct teacher distributions derived from feedback: one encourages preserving task-relevant information for utility, while the other enforces minimal and appropriate disclosure. This complementary formulation induces a Product-of-Experts (PoE) target, aligning the policy with the intersection of capability and privacy requirements. Empirical evaluations demonstrate that SELFCI, without relying on costly external supervision, consistently outperforms competitive baselines such as online reinforcement learning algorithms (e.g., GRPO). These trends further extend to out-of-domain settings involving agentic workflows and accumulated private context, suggesting that SELFCI provides a practical path toward CI alignment.

隐私对齐大模型自蒸馏智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。