arXiv:2605.24765cs.CRcs.LG2026-05

构建隐私保护的网络安全问答基准,评估大模型在真实场景中的安全与隐私平衡能力。

CyberMaskQA: A Privacy-Aware Benchmark for Evaluating Large Language Models in Cybersecurity Question Answering

论文配图:CyberMaskQA: A Privacy-Aware Benchmark for Evaluating Large Language Models in Cybersecurity Question Answering
图 1 · 摘自论文原文
  • 基于真实组织上下文生成带敏感信息标注的问答数据集
  • 支持在保留隐私前提下进行精准安全推理,准确率与掩码效果可量化评估
  • 适合研究隐私保护模型、安全AI部署及合规性验证的开发者和研究人员

大型语言模型(LLMs)正被广泛应用于网络安全问答(QA),用于事件响应与漏洞分析等关键任务。然而,实际操作中涉及系统日志、网络配置等敏感信息,如IP地址、主机名、用户账号。在受监管环境中使用云端模型处理此类数据存在安全隐患或不可行。此外,隐私保护型问答的发展受限于缺乏兼具标注丰富性与上下文真实性的数据集。为此,我们提出CYBERMASKQA——一个面向网络安全核心领域的隐私感知问答基准。不同于仅测试事实知识的现有基准,该数据集基于真实组织上下文,明确刻画资产与权限间的因果依赖关系。通过系统化流程生成,结合人工设计基础场景与大模型语义扩展,并为每条实例精确标注私有实体,支持可控信息披露。对问答准确率与遮蔽性能的评估表明,该基准可用于开发可部署、上下文感知的网络安全模型,并推动对隐私-效用权衡的深入研究。论文接受后将公开数据集与生成框架。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly applied to cybersecurity question answering (QA) for critical tasks such as incident response and vulnerability analysis. However, real-world operational contexts, including system logs and network configurations, inherently contain sensitive identifiers, e.g., IP addresses, host names, and user accounts. Processing this data with cloud-based models is often unsafe or infeasible in regulated environments. Furthermore, progress in privacy-preserving QA is hindered by the lack of annotated, context-rich datasets capable of jointly evaluating operational reasoning and privacy preservation. To address this gap, we introduce CYBERMASKQA, a privacy-aware QA benchmark covering key security domains. Unlike existing benchmarks that primarily test factual knowledge, CYBERMASKQA grounds questions in realistic organizational contexts with explicit causal dependencies among assets and privileges. Generated through a systematic pipeline, the dataset combines human-curated base scenarios with LLM-driven semantic expansion, annotating each instance with precise private entity labels to enable controlled information disclosure. Evaluations of QA accuracy and masking performance demonstrate the benchmark's utility for developing deployable, context-aware cybersecurity models and facilitating nuanced studies of privacy-utility trade-offs. Upon acceptance, we will release the dataset and the generation framework.

网络安全隐私保护大模型评测问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。