大模型能从聊天记录推断用户性格,存在隐私泄露风险
Can LLMs Infer Conversational Agent Users' Personality Traits from Chat History?
- 用RoBERTa模型分析聊天记录,预测人格特质
- 外向性预测准确率比基线高44%(关系与反思类对话)
- 提醒用户注意聊天内容可能暴露性格信息
敏感信息如个人性格特征可能被用于操纵行为(如个性化推送)。为评估基于大语言模型的对话代理(CAs)中用户交互数据的隐私风险,我们收集了668名参与者的真实ChatGPT日志,包含62,090条对话记录,并报告了不同类型共享数据及使用场景的统计结果。我们微调了RoBERTa-base文本分类模型,以从CA交互中推断人格特质。结果显示,该模型在多项人格特质的三分类任务中表现优于随机猜测。例如,在涉及人际关系和个人反思的对话中,外向性预测准确率相对基线提升44%。研究揭示了与CAs互动带来的隐私风险,并提供了不同交互类型风险水平的精细化洞察。
原文摘要 · Abstract (English)
Sensitive information, such as knowledge about an individual's personality, can be can be misused to influence behavior (e.g., via personalized messaging). To assess to what extent an individual's personality can be inferred from user interactions with LLM-based conversational agents (CAs), we analyze and quantify related privacy risks of using CAs. We collected actual ChatGPT logs from N=668 participants, containing 62,090 individual chats, and report statistics about the different types of shared data and use cases. We fine-tuned RoBERTa-base text classification models to infer personality traits from CA interactions. The findings show that these models achieve trait inference with accuracy (ternary classification) better than random in multiple cases. For example, for extraversion, accuracy improves by +44% relative to the baseline on interactions for relationships and personal reflection. This research highlights how interactions with CAs pose privacy risks and provides fine-grained insights into the level of risk associated with different types of interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。