用大模型代理生成逼真数字足迹,解决数据稀缺难题。
PersonaTrace: Synthesizing Realistic Digital Footprints with LLM Agents
- 基于用户画像,用大模型代理生成多样且可信的数字行为序列。
- 生成数据在多样性与真实性上超越现有基线,提升下游任务表现。
- 适合研究个性化应用、行为建模及隐私保护技术的学者使用。
数字足迹(个体与数字系统的交互记录)对行为研究、个性化应用开发及机器学习模型训练至关重要。然而,该领域常受限于数据多样性不足与获取困难。为此,我们提出一种新方法,利用大语言模型(LLM)代理合成逼真的数字足迹。从结构化用户画像出发,该方法生成多样且合理的用户事件序列,最终产出电子邮件、消息、日历条目、提醒等数字产物。内在评估显示,生成数据在多样性和真实性上优于现有基线。此外,在真实世界分布外任务上,基于该合成数据微调的模型性能超过其他合成数据训练的模型。
原文摘要 · Abstract (English)
Digital footprints (records of individuals' interactions with digital systems) are essential for studying behavior, developing personalized applications, and training machine learning models. However, research in this area is often hindered by the scarcity of diverse and accessible data. To address this limitation, we propose a novel method for synthesizing realistic digital footprints using large language model (LLM) agents. Starting from a structured user profile, our approach generates diverse and plausible sequences of user events, ultimately producing corresponding digital artifacts such as emails, messages, calendar entries, reminders, etc. Intrinsic evaluation results demonstrate that the generated dataset is more diverse and realistic than existing baselines. Moreover, models fine-tuned on our synthetic data outperform those trained on other synthetic datasets when evaluated on real-world out-of-distribution tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。