测试文化敏感查询是否导致RAG系统泄露更多个人隐私,发现无显著放大效应。
Exploratory As-Analyzed No-Detection of Culturally-Marked Predicate-Triggered PII Amplification in a Synthetic-English RAG Probe: A Predicate-Resource-Confounded Audit
- 设计四文化对比实验,用合成数据测试刻板印象触发的隐私泄露
- 经多重检验后,四种文化下均未发现刻板印象引发的隐私泄露放大
- 适用于关注AI伦理与隐私风险审计的研究者
我们探讨针对文化标记人群的刻板印象型查询,是否会比同等条件下的中性查询从检索增强生成(RAG)系统中泄露更多个人信息。在合成英文个人身份信息(PII)语料上,预注册了包含四个文化群体(en-Anglo、es-LATAM、阿拉伯语、印地语)的审计,比较五种查询类型,称为刻板印象触发泄漏差值(STLD)。需强调两点:其一,确认性估计量从未运行,本文所有测试均为探索性或敏感性分析,计划偏差均列于附录;其二,名称泄露指标受提示回声伪影污染:模型常直接重述所问姓名,造成虚假泄露膨胀且不依赖真实检索。在更清洁的通道(邮箱、电话、社保号类、地址)中,经过多重比较校正后,四种文化均未发现刻板印象驱动的隐私泄露放大。由于样本仅对中等效应有统计效力,且文化标记探针混合了刻板印象内容与文化标识及传统实践,本研究报告为‘未检测到’,而非‘无影响’,尤其针对资源混杂背景下的文化标记谓词泄露。
原文摘要 · Abstract (English)
We ask whether stereotype-loaded queries about culturally marked people leak more personal information from a retrieval-augmented generation (RAG) system than otherwise-equivalent neutral queries. We pre-register a four-culture audit (en-Anglo, es-LATAM, Arabic, Hindi) on a synthetic English PII corpus, comparing five query arms we call the Stereotype-Trigger Leakage Delta (STLD). Two caveats up front. Our locked confirmatory estimator was never run, so every test in the paper is exploratory or sensitivity, with all plan deviations listed in the appendix. And the name-leakage metric is contaminated by a prompt-echo artifact: the model often just re-emits the name we asked about, which inflates apparent leakage without any retrieval at all. On the cleaner channels (email, phone, ssn-like, address), we find no stereotype-driven amplification on any of the four cultures after multiple-comparison correction. Because our sample is only powered for mid-sized effects, and because the culturally marked probes mix stereotype content with cultural markers and heritage practices, we present this as no detection, not evidence of no effect, of culturally marked predicate leakage that is confounded with the underlying resource.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。