arXiv:2501.19298cs.AIcs.LG2025-01被引 5

用大模型生成智能家居行为数据,解决真实数据难收集、过时问题。

Synthetic User Behavior Sequence Generation with Large Language Models for Smart Homes

  • 基于大模型设计结构感知压缩方法,高效生成行为序列。
  • 自动生成符合规范的合成数据,提升下游模型泛化能力。
  • 适合智能安防、行为预测等需动态数据的场景研究者。

近年来,随着智能家居系统广泛应用,其安全威胁日益突出。现有安全方案多依赖预先收集的固定数据集进行训练,但数据采集耗时且难以适应不断变化的环境,同时涉及用户隐私风险。大语言模型(LLMs)在自然语言处理、推理和问题解决方面展现出强大能力。本文提出IoTGen框架,利用大模型生成反映环境变化的合成数据,以增强下游智能模型的泛化性能。首先,提出面向物联网行为数据的结构模式感知压缩(SPPC)方法,有效减少令牌消耗并保留关键信息;其次,构建系统化提示工程与数据生成流程,自动产生具有合理性和规范性的合成物联网数据,支持任务模型自适应训练,从而提升在真实场景中的表现。

原文摘要 · Abstract (English)

In recent years, as smart home systems have become more widespread, security concerns within these environments have become a growing threat. Currently, most smart home security solutions, such as anomaly detection and behavior prediction models, are trained using fixed datasets that are precollected. However, the process of dataset collection is time-consuming and lacks the flexibility needed to adapt to the constantly evolving smart home environment. Additionally, the collection of personal data raises significant privacy concerns for users. Lately, large language models (LLMs) have emerged as a powerful tool for a wide range of tasks across diverse application domains, thanks to their strong capabilities in natural language processing, reasoning, and problem-solving. In this paper, we propose an LLM-based synthetic dataset generation IoTGen framework to enhance the generalization of downstream smart home intelligent models. By generating new synthetic datasets that reflect changes in the environment, smart home intelligent models can be retrained to overcome the limitations of fixed and outdated data, allowing them to better align with the dynamic nature of real-world home environments. Specifically, we first propose a Structure Pattern Perception Compression (SPPC) method tailored for IoT behavior data, which preserves the most informative content in the data while significantly reducing token consumption. Then, we propose a systematic approach to create prompts and implement data generation to automatically generate IoT synthetic data with normative and reasonable properties, assisting task models in adaptive training to improve generalization and real-world performance.

智能家居大模型数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。