用博弈论平衡大模型问答的隐私与准确,防止敏感信息泄露。
Beyond Local vs. External: A Game-Theoretic Framework for Trustworthy Knowledge Acquisition

- 设计对抗性子查询生成器,将敏感问题拆成低风险片段。
- 在生物医学和法律数据集上,泄露率降低60%以上,回答准确率仍高。
- 适合对隐私要求高的医疗、法律等敏感领域使用。
云端大语言模型虽具强大推理与动态知识能力,但直接提交原始查询会暴露用户敏感意图。完全依赖本地模型虽保隐私,却因参数量和知识有限而影响答案质量。为此,我们提出博弈论可信知识获取框架(GTKA),将知识效用与隐私保护建模为战略博弈。GTKA包含三部分:(i) 隐私感知的子查询生成器,将敏感意图分解为泛化、低风险片段;(ii) 对抗重构攻击者,尝试从片段中还原原始查询,提供自适应泄露信号;(iii) 可信本地整合器,在安全边界内合成外部响应。通过交替训练生成器与攻击者,优化子查询生成策略,在最大化知识获取准确率的同时最小化原始意图可重构性。我们在生物医学与法律领域构建两个敏感场景基准。大量实验表明,相比最先进基线,GTKA显著降低意图泄露,同时保持高保真回答质量。
原文摘要 · Abstract (English)
Cloud-hosted Large Language Models (LLMs) offer unmatched reasoning capabilities and dynamic knowledge, yet submitting raw queries to these external services risks exposing sensitive user intent. Conversely, relying exclusively on trusted local models preserves privacy but often compromises answer quality due to limited parameter scale and knowledge. To resolve this dilemma, we propose Game-theoretic Trustworthy Knowledge Acquisition (GTKA), a framework that formulates the trade-off between knowledge utility and privacy as a strategic game. GTKA consists of three components: (i) a privacy-aware sub-query generator that decomposes sensitive intent into generalized, low-risk fragments; (ii) an adversarial reconstruction attacker that attempts to infer the original query from these fragments, providing adaptive leakage signals; and (iii) a trusted local integrator that synthesizes external responses within a secure boundary. By training the generator and attacker in an alternating adversarial manner, GTKA optimizes the sub-query generation policy to maximize knowledge acquisition accuracy while minimizing the reconstructability of the original sensitive intent. To validate our approach, we construct two sensitive-domain benchmarks in the biomedical and legal fields. Extensive experiments demonstrate that GTKA significantly reduces intent leakage compared to state-of-the-art baselines while maintaining high-fidelity answer quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。