通过分析安卓应用隐私漏洞,提出可解释的机器学习框架提升敏感信息检测准确率。
Guarding Digital Privacy: Exploring User Profiling and Security Enhancements
- 结合决策树与神经网络,用LIME增强模型可解释性。
- 在18款印度应用上实现75.01%准确率,训练仅需3.62秒。
- 适合关注移动隐私保护与可解释AI的研究者参考。
用户画像通过收集用户信息以提供个性化推荐,已广泛应用于技术发展,但其增长对用户隐私构成威胁,因设备常在用户不知情下采集敏感数据。本文综述用户画像相关方法与挑战,基于两家公司共享用户数据及对18款印度热门安卓应用(涵盖社交、教育、娱乐、旅游、购物及其他类别)的分析,揭示了隐私漏洞。进一步提出一种改进的机器学习框架,融合决策树与神经网络,在检测个人敏感信息暴露方面优于现有分类器。利用可解释人工智能算法LIME(局部可解释模型无关解释),提升模型可解释性,有助于可靠识别敏感数据。实验结果显示,该框架在神经网络上实现75.01%的准确率,训练时间缩短至3.62秒。最后,论文提出未来研究方向以强化数字安全措施。
原文摘要 · Abstract (English)
User profiling, the practice of collecting user information for personalized recommendations, has become widespread, driving progress in technology. However, this growth poses a threat to user privacy, as devices often collect sensitive data without their owners' awareness. This article aims to consolidate knowledge on user profiling, exploring various approaches and associated challenges. Through the lens of two companies sharing user data and an analysis of 18 popular Android applications in India across various categories, including $\textit{Social, Education, Entertainment, Travel, Shopping and Others}$, the article unveils privacy vulnerabilities. Further, the article propose an enhanced machine learning framework, employing decision trees and neural networks, that improves state-of-the-art classifiers in detecting personal information exposure. Leveraging the XAI (explainable artificial intelligence) algorithm LIME (Local Interpretable Model-agnostic Explanations), it enhances interpretability, crucial for reliably identifying sensitive data. Results demonstrate a noteworthy performance boost, achieving a $75.01\%$ accuracy with a reduced training time of $3.62$ seconds for neural networks. Concluding, the paper suggests research directions to strengthen digital security measures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。