arXiv:2511.15434cs.CRcs.AI2025-11被引 5

用小模型检测钓鱼网站,兼顾成本、性能与隐私。

Small Language Models for Phishing Website Detection: Cost, Performance, and Privacy Trade-Offs

  • 仅用网页原始HTML代码,测试15个参数量10亿到700亿的小语言模型。
  • 小模型在准确率上略逊于大模型,但可本地部署且成本更低。
  • 适合注重数据隐私和预算有限的组织使用。

钓鱼网站是重大网络安全威胁,常导致严重财务与组织损失。传统机器学习方法需大量特征工程、持续重训及高昂运维成本。虽然大型语言模型(LLMs)在钓鱼网站分类任务中表现优异,但其运营成本高且依赖外部服务,限制了实际应用。本文研究仅用原始HTML代码检测钓鱼网站的轻量级语言模型(SLMs)可行性。评估了15个参数量介于10亿至700亿之间的常见SLMs,对比其分类准确率、计算需求与成本效益。结果表明,尽管SLMs性能不及顶尖商用LLMs,但仍能提供可替代外部服务的可行、可扩展方案。通过成本与收益的对比分析,本研究为未来在钓鱼检测系统中适配、微调与部署SLMs奠定基础,力求在安全效果与经济实用性间取得平衡。

原文摘要 · Abstract (English)

Phishing websites pose a major cybersecurity threat, exploiting unsuspecting users and causing significant financial and organisational harm. Traditional machine learning approaches for phishing detection often require extensive feature engineering, continuous retraining, and costly infrastructure maintenance. At the same time, proprietary large language models (LLMs) have demonstrated strong performance in phishing-related classification tasks, but their operational costs and reliance on external providers limit their practical adoption in many business environments. This paper investigates the feasibility of small language models (SLMs) for detecting phishing websites using only their raw HTML code. A key advantage of these models is that they can be deployed on local infrastructure, providing organisations with greater control over data and operations. We systematically evaluate 15 commonly used Small Language Models (SLMs), ranging from 1 billion to 70 billion parameters, benchmarking their classification accuracy, computational requirements, and cost-efficiency. Our results highlight the trade-offs between detection performance and resource consumption, demonstrating that while SLMs underperform compared to state-of-the-art proprietary LLMs, they can still provide a viable and scalable alternative to external LLM services. By presenting a comparative analysis of costs and benefits, this work lays the foundation for future research on the adaptation, fine-tuning, and deployment of SLMs in phishing detection systems, aiming to balance security effectiveness and economic practicality.

小模型钓鱼检测隐私保护成本优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。