arXiv:2503.22406cs.CRcs.AI2025-03被引 3

用大模型识别拼写错误域名,准确率达98%。

Training Large Language Models for Advanced Typosquatting Detection

  • 用字符变换和模式规则训练大模型,不依赖特定域名数据。
  • 微调后Phi-4 14B模型仅用几千样本即达98%准确率。
  • 适合安全团队提升域名欺骗检测能力。

拼写错误劫持是一种长期存在的网络威胁,利用用户输入失误诱导其访问恶意网站、传播恶意软件或实施钓鱼攻击。随着域名数量和新顶级域(TLD)的激增,此类攻击手段日益复杂,对个人、企业及国家网络安全构成重大风险。传统检测方法多聚焦于已知仿冒模式,难以应对更复杂的攻击。本研究提出一种基于大语言模型(LLM)的新方法,通过在字符级变换和模式规则上训练模型,而非使用特定域名数据,构建更具适应性和鲁棒性的检测机制。实验表明,在适当微调后,Phi-4 14B 模型在仅使用数千个训练样本的情况下,达到了98%的准确率。该研究展示了大模型在网络安全中的潜力,特别是在防范域名误导性攻击方面,并为优化机器学习策略提供了启示。

原文摘要 · Abstract (English)

Typosquatting is a long-standing cyber threat that exploits human error in typing URLs to deceive users, distribute malware, and conduct phishing attacks. With the proliferation of domain names and new Top-Level Domains (TLDs), typosquatting techniques have grown more sophisticated, posing significant risks to individuals, businesses, and national cybersecurity infrastructure. Traditional detection methods primarily focus on well-known impersonation patterns, leaving gaps in identifying more complex attacks. This study introduces a novel approach leveraging large language models (LLMs) to enhance typosquatting detection. By training an LLM on character-level transformations and pattern-based heuristics rather than domain-specific data, a more adaptable and resilient detection mechanism develops. Experimental results indicate that the Phi-4 14B model outperformed other tested models when properly fine tuned achieving a 98% accuracy rate with only a few thousand training samples. This research highlights the potential of LLMs in cybersecurity applications, specifically in mitigating domain-based deception tactics, and provides insights into optimizing machine learning strategies for threat detection.

大模型安全检测域名防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。