arXiv:2505.11417cs.HCcs.AI2025-05被引 1

构建边缘设备用户画像数据集,让小模型也能懂用户习惯。

EdgeWisePersona: A Dataset for On-Device User Profiling from Natural Language Interactions

  • 用大模型生成多轮对话模拟真实家庭交互行为
  • 小模型重建用户习惯准确率远低于大模型
  • 适合研究隐私保护的本地化智能系统开发者

本文提出一个新型数据集与评估基准,用于测试和提升可部署在边缘设备上的小型语言模型,重点聚焦于从智能家居环境中的多会话自然语言交互中进行用户画像建模。数据集核心是结构化用户档案,每个档案由一系列例行任务构成——即由上下文触发、可重复的行为模式,决定用户如何与家居系统互动。利用这些档案作为输入,大语言模型(LLM)生成了模拟真实、多样且具备上下文感知能力的用户-设备对话会话。该数据集支持的核心任务是画像重建:仅根据历史交互记录推断用户例行任务与偏好。为评估当前模型在真实条件下的表现,我们对多种先进紧凑型语言模型进行了基准测试,并与大型基础模型进行对比。结果显示,尽管小模型具备一定画像重建能力,但在准确捕捉用户行为方面仍显著落后于大模型。这一性能差距构成重大挑战,尤其考虑到边缘计算在保护用户隐私、降低延迟、实现无需云端依赖的个性化体验方面具有关键优势。本数据集提供了一个真实、结构化的测试平台,可用于开发与评估在资源受限条件下建模用户行为的方法,是迈向能在用户自有设备上学习与自适应的智能、隐私友好型AI系统的关键一步。

原文摘要 · Abstract (English)

This paper introduces a novel dataset and evaluation benchmark designed to assess and improve small language models deployable on edge devices, with a focus on user profiling from multi-session natural language interactions in smart home environments. At the core of the dataset are structured user profiles, each defined by a set of routines - context-triggered, repeatable patterns of behavior that govern how users interact with their home systems. Using these profiles as input, a large language model (LLM) generates corresponding interaction sessions that simulate realistic, diverse, and context-aware dialogues between users and their devices. The primary task supported by this dataset is profile reconstruction: inferring user routines and preferences solely from interactions history. To assess how well current models can perform this task under realistic conditions, we benchmarked several state-of-the-art compact language models and compared their performance against large foundation models. Our results show that while small models demonstrate some capability in reconstructing profiles, they still fall significantly short of large models in accurately capturing user behavior. This performance gap poses a major challenge - particularly because on-device processing offers critical advantages, such as preserving user privacy, minimizing latency, and enabling personalized experiences without reliance on the cloud. By providing a realistic, structured testbed for developing and evaluating behavioral modeling under these constraints, our dataset represents a key step toward enabling intelligent, privacy-respecting AI systems that learn and adapt directly on user-owned devices.

用户画像边缘计算小模型对话生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。