arXiv:2605.10808cs.CRcs.AI2026-05

对比领域适配大模型在5G安全威胁建模中的表现,发现模型规模和解码策略影响显著。

Threat Modelling using Domain-Adapted Language Models: Empirical Evaluation and Insights

论文配图:Threat Modelling using Domain-Adapted Language Models: Empirical Evaluation and Insights
图 1 · 摘自论文原文
  • 用8个不同规模的领域适配模型评估STRIDE威胁分类效果
  • 大模型性能提升不持续,解码方式对输出有效性影响大
  • 提示词设计需针对STRIDE框架优化,避免无效输出

大型语言模型(LLMs)在网络安全应用中日益受到关注,如漏洞检测。在威胁建模领域,以往研究主要在有限提示设置下评估通用大模型。本研究通过系统评估不同规模的领域适配语言模型(涵盖电信与网络安全领域)在结构化威胁建模中的表现,扩展了该研究方向。采用广泛使用的STRIDE方法,聚焦5G安全场景。共进行52种配置实验(基于8个语言模型),分析领域适配、模型规模、解码策略(贪婪与随机采样)及提示技术对STRIDE威胁分类的影响。结果表明:领域适配模型并非始终优于通用模型;解码策略显著影响模型行为与输出有效性;虽然大模型整体表现更好,但收益并不稳定且不足以支撑可靠威胁建模。这些发现揭示当前大模型在结构化威胁建模任务中的根本局限,表明仅靠增加数据或扩大模型规模无法解决问题,亟需引入更任务特定的推理机制和更强的安全概念锚定。本文还总结了常见无效输出类型,并提出针对STRIDE建模的定制化提示建议。

原文摘要 · Abstract (English)

Large Language Models(LLMs) are increasingly explored for cybersecurity applications such as vulnerability detection. In the domain of threat modelling, prior work has primarily evaluated a number of general-purpose Large Language Models under limited prompting settings. In this study, we extend the research area of structured threat modelling by systematically evaluating domain-adapted language models of different sizes to their general counterparts. We use both LLMs and Small Language Models(SLMs) that were domain adapted to telecommunications and cybersecuirty. For the structured threat modelling, we selected the widely used STRIDE approach and the application area is 5G security. We present a comprehensive empirical evaluation using 52 different configurations (on 8 different language models) to analyze the impact of 1) domain adaptation, 2) model scale, 3) decoding strategies (greedy vs. stochastic sampling), and 4) prompting technique on STRIDE threat classification. Our results show that domain-adapted models do not consistently outperform their general-purpose counterparts, and decoding strategies significantly affect model behavior and output validity. They also show that while larger models generally achieve higher performance, these gains are neither consistent nor sufficient for reliable threat modelling. These findings highlight fundamental limitations of current LLMs for structured threat modelling tasks and suggest that improvements require more than additional training data or model scaling, motivating the need for incorporating more task-specific reasoning and stronger grounding in security concepts. We present insights on invalid outputs encountered and present suggestions for prompting tailored specifically for STRIDE threat modelling.

威胁建模大模型5G安全提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。