首次大规模分析大模型应用冒名与克隆问题
LLM App Squatting and Cloning
- 构建工具LLMappCrazy,结合相似度与语义分析识别克隆应用
- 发现超5000个冒名应用,9575起克隆案例,近20%具恶意行为
- 适用于平台安全团队、开发者及对大模型生态安全关注者
冒名和克隆攻击在移动应用商店中长期存在,恶意者利用热门应用的名称与声誉欺骗用户。随着大语言模型(LLM)应用商店如GPT Store和FlowGPT的兴起,此类问题同样浮现,威胁LLM应用生态的完整性。本研究首次基于自研工具LLMappCrazy,开展大规模分析。该工具涵盖14种冒名生成技术,并集成莱文斯坦距离与BERT语义分析,通过功能相似性检测克隆应用。我们对前1000个应用名称生成变体,发现数据集中有超过5000个冒名应用;在六个主要平台共观测到3509个冒名应用和9575个克隆案例。抽样分析显示,18.7%的冒名应用和4.9%的克隆应用存在恶意行为,包括钓鱼、恶意软件分发、虚假内容传播及激进广告注入。
原文摘要 · Abstract (English)
Impersonation tactics, such as app squatting and app cloning, have posed longstanding challenges in mobile app stores, where malicious actors exploit the names and reputations of popular apps to deceive users. With the rapid growth of Large Language Model (LLM) stores like GPT Store and FlowGPT, these issues have similarly surfaced, threatening the integrity of the LLM app ecosystem. In this study, we present the first large-scale analysis of LLM app squatting and cloning using our custom-built tool, LLMappCrazy. LLMappCrazy covers 14 squatting generation techniques and integrates Levenshtein distance and BERT-based semantic analysis to detect cloning by analyzing app functional similarities. Using this tool, we generated variations of the top 1000 app names and found over 5,000 squatting apps in the dataset. Additionally, we observed 3,509 squatting apps and 9,575 cloning cases across six major platforms. After sampling, we find that 18.7% of the squatting apps and 4.9% of the cloning apps exhibited malicious behavior, including phishing, malware distribution, fake content dissemination, and aggressive ad injection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。