arXiv:2504.04332cs.CLcs.AI2025-04被引 10

小模型也能高仿真人写作,引出隐私与安全新风险

IMPersona: Evaluating Individual Level LM Impersonation

  • 用微调+记忆检索让小模型模仿特定人写作风格
  • 盲测中44.44%对话被误认成真人,远超基线方法
  • 为防伪造提供检测与防御方案,适合关注安全的开发者

随着语言模型在对话生成中愈发接近人类表现,一个关键问题浮现:它们能在多大程度上模拟特定个体的特征?为此,我们提出IMPersona框架,用于评估语言模型在模仿个人写作风格和私人知识方面的表现。通过监督微调与受层次记忆启发的检索系统,我们发现即便是规模较小的开源模型(如Llama-3.1-8B-Instruct),也能达到令人担忧的模仿水平。在盲测对话实验中,集成记忆的微调模型有44.44%的交互被参与者误认为是真人,而最佳提示方法仅为25.00%。我们分析结果并提出相应的检测方法与防御策略。研究揭示了个性化语言模型在应用中的潜力与风险,尤其涉及隐私、安全及伦理部署等现实挑战。

原文摘要 · Abstract (English)

As language models achieve increasingly human-like capabilities in conversational text generation, a critical question emerges: to what extent can these systems simulate the characteristics of specific individuals? To evaluate this, we introduce IMPersona, a framework for evaluating LMs at impersonating specific individuals' writing style and personal knowledge. Using supervised fine-tuning and a hierarchical memory-inspired retrieval system, we demonstrate that even modestly sized open-source models, such as Llama-3.1-8B-Instruct, can achieve impersonation abilities at concerning levels. In blind conversation experiments, participants (mis)identified our fine-tuned models with memory integration as human in 44.44% of interactions, compared to just 25.00% for the best prompting-based approach. We analyze these results to propose detection methods and defense strategies against such impersonation attempts. Our findings raise important questions about both the potential applications and risks of personalized language models, particularly regarding privacy, security, and the ethical deployment of such technologies in real-world contexts.

语言模型个性模仿安全风险隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。