arXiv:2409.03735cs.LGcs.AI2024-09被引 5

用情境完整性评估大模型隐私偏见,判断信息泄露风险。

Privacy Bias in Language Models: A Contextual Integrity-based Auditing Metric

  • 基于情境完整性构建隐私偏见评估方法,量化信息流是否恰当。
  • 发现模型能力与优化策略显著影响隐私偏见程度。
  • 适合模型训练者、服务提供方和政策制定者做伦理审查。

随着大型语言模型(LLMs)被整合进社会技术系统,评估其表现出的隐私偏见至关重要。我们定义隐私偏见为模型响应中信息流的适当性价值。隐私偏见与预期值之间的偏差(称为隐私偏见差)可能指示隐私违规。作为审计指标,隐私偏见可帮助(a)模型训练者评估LLMs的伦理与社会影响,(b)服务提供商选择上下文合适的LLMs,(c)政策制定者评估部署中LLMs的隐私偏见合理性。我们提出并回答了一个新研究问题:如何可靠地检测LLMs中的隐私偏见及其影响因素?我们提出一种基于情境完整性的方法,评估不同LLMs的响应。该方法考虑了提示变化下响应敏感性的差异,这会阻碍隐私偏见评估。最后,我们研究了模型能力与优化对隐私偏见的影响。

原文摘要 · Abstract (English)

As large language models (LLMs) are integrated into sociotechnical systems, it is crucial to examine the privacy biases they exhibit. We define privacy bias as the appropriateness value of information flows in responses from LLMs. A deviation between privacy biases and expected values, referred to as privacy bias delta, may indicate privacy violations. As an auditing metric, privacy bias can help (a) model trainers evaluate the ethical and societal impact of LLMs, (b) service providers select context-appropriate LLMs, and (c) policymakers assess the appropriateness of privacy biases in deployed LLMs. We formulate and answer a novel research question: how can we reliably examine privacy biases in LLMs and the factors that influence them? We present a novel approach for assessing privacy biases using a contextual integrity-based methodology to evaluate the responses from various LLMs. Our approach accounts for the sensitivity of responses across prompt variations, which hinders the evaluation of privacy biases. Finally, we investigate how privacy biases are affected by model capacities and optimizations.

隐私保护大模型评估情境完整性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。