arXiv:2601.10524cs.AI2026-01

对比多个大模型在钓鱼检测中的泛化能力,发现架构与数据多样性决定模型可靠性。

Diagnosing Generalization Failures in Fine-Tuned LLMs: A Cross-Architectural Study on Phishing Detection

  • 跨架构对比3个模型,用SHAP和可解释性分析失败原因
  • Gemma 2 9B在多样化数据上达91%以上F1,Llama 3.1 8B无法融合多源数据
  • Mistral模型表现稳定,适合需要高鲁棒性的实际应用

微调大语言模型(LLMs)在特定任务上已达到顶尖性能,但诊断其为何变得脆弱、难以泛化仍是关键难题。为此,我们引入并应用多层次诊断框架,开展跨架构研究。将Llama 3.1 8B、Gemma 2 9B和Mistral模型在高风险钓鱼检测任务上进行微调,并通过SHAP分析与机制可解释性揭示泛化失败的根本原因。研究发现:(1)泛化依赖架构与数据多样性的协同作用;仅当Gemma 2 9B在风格多样的“通用”数据集上训练时,性能达>91% F1;(2)泛化高度依赖架构;我们诊断出Llama 3.1 8B在窄域表现良好,却无法整合多样化数据,导致显著性能下降;(3)部分架构天生更具泛化能力;Mistral模型在多种训练范式下均表现一致且稳健。本工作为诊断和理解泛化失败提供具体方法,强调可靠AI需深入验证架构、数据与训练策略的相互作用。

原文摘要 · Abstract (English)

The practice of fine-tuning Large Language Models (LLMs) has achieved state-of-the-art performance on specialized tasks, yet diagnosing why these models become brittle and fail to generalize remains a critical open problem. To address this, we introduce and apply a multi-layered diagnostic framework to a cross-architectural study. We fine-tune Llama 3.1 8B, Gemma 2 9B, and Mistral models on a high-stakes phishing detection task and use SHAP analysis and mechanistic interpretability to uncover the root causes of their generalization failures. Our investigation reveals three critical findings: (1) Generalization is driven by a powerful synergy between architecture and data diversity. The Gemma 2 9B model achieves state-of-the-art performance (>91\% F1), but only when trained on a stylistically diverse ``generalist'' dataset. (2) Generalization is highly architecture-dependent. We diagnose a specific failure mode in Llama 3.1 8B, which performs well on a narrow domain but cannot integrate diverse data, leading to a significant performance drop. (3) Some architectures are inherently more generalizable. The Mistral model proves to be a consistent and resilient performer across multiple training paradigms. By pinpointing the flawed heuristics responsible for these failures, our work provides a concrete methodology for diagnosing and understanding generalization failures, underscoring that reliable AI requires deep validation of the interplay between architecture, data, and training strategy.

大模型泛化能力可解释性钓鱼检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。