用大模型自动简化漏洞描述,让非专业人士也能看懂。
Automatic Simplification of Common Vulnerabilities and Exposures Descriptions
- 用大模型直接简化漏洞文本,不需人工标注
- 简化后文本虽更易读但常丢失关键信息
- 适合安全科普、初学者或非技术决策者阅读
网络安全对个人和组织日益重要,但相关术语对非专业人士难以理解。本文研究大语言模型在自动文本简化(ATS)中的应用,聚焦于常见漏洞披露(CVE)描述的可读性提升。目前该领域尚未有系统研究。研究构建了首个网络安全领域的ATS基准和包含40个CVE描述的测试集,并通过两轮专家问卷评估。结果显示,尽管现成大模型使文本看起来更简单,但在保持原意方面表现不佳。代码与数据已公开。
原文摘要 · Abstract (English)
Understanding cyber security is increasingly important for individuals and organizations. However, a lot of information related to cyber security can be difficult to understand to those not familiar with the topic. In this study, we focus on investigating how large language models (LLMs) could be utilized in automatic text simplification (ATS) of Common Vulnerability and Exposure (CVE) descriptions. Automatic text simplification has been studied in several contexts, such as medical, scientific, and news texts, but it has not yet been studied to simplify texts in the rapidly changing and complex domain of cyber security. We created a baseline for cyber security ATS and a test dataset of 40 CVE descriptions, evaluated by two groups of cyber security experts in two survey rounds. We have found that while out-of-the box LLMs can make the text appear simpler, they struggle with meaning preservation. Code and data are available at https://version.aalto.fi/gitlab/vehomav1/simplification\_nmi.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。