arXiv:2509.13892cs.HCcs.AI2025-09

用大模型生成手机使用数据,解决真实数据难收集问题。

Synthetic Data Generation for Screen Time and App Usage

  • 用四种提示策略生成合成手机使用数据,测试不同设计效果。
  • 详细提示+真实样例能提升数据的结构化与行为合理性。
  • 适合需要隐私保护或快速获取数据的研究者使用。

智能手机使用数据可为理解人机交互与人类行为提供重要洞察。但大规模真实场景数据采集面临高成本、隐私顾虑、样本代表性不足及非响应偏差等问题,影响结果可靠性。本文探索利用大语言模型(如OpenAI的ChatGPT)生成合成手机使用数据,作为替代方案。通过案例研究比较四种提示策略对生成数据质量的影响,重点分析提示细节程度(用户画像描述、结果特征描述)与是否引入初始真实样例两个因素。结果表明,在使用详细提示时,生成数据具备良好的结构与行为合理性,适用于部分研究场景。然而,单一合成数据集仍难以涵盖人类行为的多样性,需在数据保真度与多样性间权衡,建议根据具体应用场景设计评估指标,并未来结合更多样种子数据与不同大模型进行研究。

原文摘要 · Abstract (English)

Smartphone usage data can provide valuable insights for understanding interaction with technology and human behavior. However, collecting large-scale, in-the-wild smartphone usage logs is challenging due to high costs, privacy concerns, under representative user samples and biases like non-response that can skew results. These challenges call for exploring alternative approaches to obtain smartphone usage datasets. In this context, large language models (LLMs) such as Open AI's ChatGPT present a novel approach for synthetic smartphone usage data generation, addressing limitations of real-world data collection. We describe a case study on how four prompt strategies influenced the quality of generated smartphone usage data. We contribute with insights on prompt design and measures of data quality, reporting a prompting strategy comparison combining two factors, prompt level of detail (describing a user persona, describing the expected results characteristics) and seed data inclusion (with versus without an initial real usage example). Our findings suggest that using LLMs to generate structured and behaviorally plausible smartphone use datasets is feasible for some use cases, especially when using detailed prompts. Challenges remain in capturing diverse nuances of human behavioral patterns in a single synthetic dataset, and evaluating tradeoffs between data fidelity and diversity, suggesting the need for use-case-specific evaluation metrics and future research with more diverse seed data and different LLM models.

合成数据大模型行为分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。