揭秘ChatGPT如何自动构建用户画像,揭示隐私风险与防护方案
The Algorithmic Self-Portrait: Deconstructing Memory in ChatGPT
- 系统自动创建96%记忆,用户缺乏控制权
- 52%记忆含心理洞察,28%含受保护个人信息
- 提出防护框架,实时提醒并优化敏感提问
为实现个性化和上下文感知交互,对话式AI系统引入了记忆机制。该机制生成我们称之为“算法自画像”的新型个性化形式,基于用户在私密对话中主动披露的信息。尽管记忆可提升对话连贯性,但其创建过程仍不透明,引发关于数据敏感性、用户自主权及画像准确性的关切。我们分析了来自80名真实用户、共计2,050条记忆条目,发现:(1)96%的记忆由系统单方面创建,可能削弱用户控制力;(2)28%的记忆包含受GDPR定义的个人数据,52%包含对用户的心理洞察;(3)84%的记忆直接源自用户上下文,表明对话内容被忠实记录。最后,我们提出“归属盾”框架,可预测潜在敏感推断,警示高风险记忆生成,并建议查询改写方式,在不牺牲功能的前提下保护隐私。
原文摘要 · Abstract (English)
To enable personalized and context-aware interactions, conversational AI systems have introduced a new mechanism: Memory. Memory creates what we refer to as the Algorithmic Self-portrait - a new form of personalization derived from users' self-disclosed information divulged within private conversations. While memory enables more coherent exchanges, the underlying processes of memory creation remain opaque, raising critical questions about data sensitivity, user agency, and the fidelity of the resulting portrait. To bridge this research gap, we analyze 2,050 memory entries from 80 real-world ChatGPT users. Our analyses reveal three key findings: (1) A striking 96% of memories in our dataset are created unilaterally by the conversational system, potentially shifting agency away from the user; (2) Memories, in our dataset, contain a rich mix of GDPR-defined personal data (in 28% memories) along with psychological insights about participants (in 52% memories); and (3)~A significant majority of the memories (84%) are directly grounded in user context, indicating faithful representation of the conversations. Finally, we introduce a framework-Attribution Shield-that anticipates these inferences, alerts about potentially sensitive memory inferences, and suggests query reformulations to protect personal information without sacrificing utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。