提出传感器数据流的算法化数据最小化方法,兼顾隐私与模型精度。
Algorithmic Data Minimization for Machine Learning over Internet-of-Things Data Streams
- 基于弱信号特性设计可量化数据最小化框架
- 降低用户可识别性16.7%,精度损失低于1%
- 适合部署在家庭、办公等敏感场景的IoT系统
机器学习可分析物联网设备生成的海量数据,以识别模式、做出预测并实现实时决策。通过处理传感器数据,机器学习模型能优化流程、提升效率并增强个性化体验。然而,物联网系统常部署于家庭、办公室等敏感环境,可能无意中暴露位置、习惯和个人标识信息,引发重大隐私担忧。这要求应用数据最小化——新兴数据法规中的基本原则,即服务提供商仅收集与特定目的直接相关且必要的数据。尽管重要,数据最小化在传感器数据背景下缺乏精确的技术定义,因弱信号集合难以适用二元的“相关必要”判断标准。本文为传感器数据流中的数据最小化提供技术解读,探索可行的实施方法并解决相关挑战。通过该方法,框架可将用户可识别性降低16.7%,同时保持精度损失低于1%,为隐私保护的物联网数据处理提供可行路径。
原文摘要 · Abstract (English)
Machine learning can analyze vast amounts of data generated by IoT devices to identify patterns, make predictions, and enable real-time decision-making. By processing sensor data, machine learning models can optimize processes, improve efficiency, and enhance personalized user experiences in smart systems. However, IoT systems are often deployed in sensitive environments such as households and offices, where they may inadvertently expose identifiable information, including location, habits, and personal identifiers. This raises significant privacy concerns, necessitating the application of data minimization -- a foundational principle in emerging data regulations, which mandates that service providers only collect data that is directly relevant and necessary for a specified purpose. Despite its importance, data minimization lacks a precise technical definition in the context of sensor data, where collections of weak signals make it challenging to apply a binary "relevant and necessary" rule. This paper provides a technical interpretation of data minimization in the context of sensor streams, explores practical methods for implementation, and addresses the challenges involved. Through our approach, we demonstrate that our framework can reduce user identifiability by up to 16.7% while maintaining accuracy loss below 1%, offering a viable path toward privacy-preserving IoT data processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。