arXiv:2505.17615cs.LGcs.CL2025-05

用大模型生成用户行为数据,既保护隐私又提升预测效果

Large language model as user daily behavior data generator: balancing population diversity and individual personality

  • 用大模型根据用户画像和真实事件模拟行为
  • 在移动和手机使用预测上最高提升18.9%
  • 适合需要隐私保护的数据增强场景

预测人类日常行为因习惯模式复杂及短期波动而困难。尽管数据驱动模型借助多平台、多设备的实证数据提升了预测能力,但对敏感大规模用户数据的依赖引发隐私担忧并限制数据获取。合成数据生成成为有前景的解决方案,但现有方法常局限于特定应用。本文提出BehaviorGen框架,利用大语言模型(LLMs)生成高质量合成行为数据。通过基于用户画像和真实事件模拟行为,BehaviorGen支持行为预测模型中的数据增强与替换。我们在多个场景下评估其性能,包括数据增强、微调替换和微调增强,显著提升了人类移动性和智能手机使用预测表现,最高提升达18.9%。结果表明,BehaviorGen可通过灵活且隐私友好的合成数据生成,增强用户行为建模能力。

原文摘要 · Abstract (English)

Predicting human daily behavior is challenging due to the complexity of routine patterns and short-term fluctuations. While data-driven models have improved behavior prediction by leveraging empirical data from various platforms and devices, the reliance on sensitive, large-scale user data raises privacy concerns and limits data availability. Synthetic data generation has emerged as a promising solution, though existing methods are often limited to specific applications. In this work, we introduce BehaviorGen, a framework that uses large language models (LLMs) to generate high-quality synthetic behavior data. By simulating user behavior based on profiles and real events, BehaviorGen supports data augmentation and replacement in behavior prediction models. We evaluate its performance in scenarios such as pertaining augmentation, fine-tuning replacement, and fine-tuning augmentation, achieving significant improvements in human mobility and smartphone usage predictions, with gains of up to 18.9%. Our results demonstrate the potential of BehaviorGen to enhance user behavior modeling through flexible and privacy-preserving synthetic data generation.

行为预测合成数据大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。