arXiv:2603.01289cs.CL2026-03被引 2

用十年聊天记录测试大模型模拟个人,发现当前方法仍难骗过熟人。

Individual Turing Test: A Case Study of LLM-based Simulation Using Longitudinal Personal Data

  • 构建‘个体图灵测试’,用真实对话数据评估大模型对特定个体的模仿能力。
  • 微调提升日常语言风格还原度,但检索与记忆增强法在观点类问题上表现更优。
  • 揭示参数化与非参数化方法在长期个性模拟中的根本性权衡,适合研究个性化生成者关注。

大型语言模型(LLMs)展现出惊人的类人能力,但其对特定个体的复现能力仍待探索。本文通过一位志愿者提供的十年私人消息历史档案,开展案例研究,探讨基于LLM的个体仿真。基于这些数据,提出“个体图灵测试”,评估熟人能否从多个候选回复中准确识别出最可能来自该志愿者的回答。研究对比了主流仿真方法:微调、检索增强生成(RAG)、基于记忆的方法及融合微调与RAG或记忆的混合方法。实证结果显示,当前方法尚未通过个体图灵测试,但在陌生人参与的测试中表现显著更优。此外,微调有助于还原日常对话的语言风格,而检索增强和基于记忆的方法在涉及个人观点与偏好的问题上表现更强。这些发现揭示了在纵向上下文下,基于参数与非参数方法进行个体仿真时的根本性权衡。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable human-like capabilities, yet their ability to replicate a specific individual remains under-explored. This paper presents a case study to investigate LLM-based individual simulation with a volunteer-contributed archive of private messaging history spanning over ten years. Based on the messaging data, we propose the "Individual Turing Test" to evaluate whether acquaintances of the volunteer can correctly identify which response in a multi-candidate pool most plausibly comes from the volunteer. We investigate prevalent LLM-based individual simulation approaches including: fine-tuning, retrieval-augmented generation (RAG), memory-based approach, and hybrid methods that integrate fine-tuning and RAG or memory. Empirical results show that current LLM-based simulation methods do not pass the Individual Turing Test, but they perform substantially better when the same test is conducted on strangers to the target individual. Additionally, while fine-tuning improves the simulation in daily chats representing the language style of the individual, retrieval-augmented and memory-based approaches demonstrate stronger performance on questions involving personal opinions and preferences. These findings reveal a fundamental trade-off between parametric and non-parametric approaches to individual simulation with LLMs when given a longitudinal context.

大模型个性模拟图灵测试长时序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。