arXiv:2502.07045cs.CRcs.AI2025-02被引 6

用大模型生成数据检测内鬼威胁,解决隐私与数据获取难题

Scalable and Ethical Insider Threat Detection through Data Synthesis and Analysis by LLMs

  • 用大模型合成匿名求职网站评论数据,避免真实数据采集伦理问题
  • 大模型情感分析结果与人工评分高度一致,能识别潜在威胁信号
  • 适合安全研究者、企业风控团队用于低成本、大规模威胁筛查

内鬼威胁虽人数少,却因内部权限造成巨大影响,例如匿名求职网站评论可能暴露组织风险。本研究探索大语言模型(LLMs)在分析此类评论中识别内鬼情绪的潜力。为应对数据采集的伦理挑战,研究结合现有评论数据集与大模型合成数据,对比了模型生成的情感分数与专家人工评分的一致性。结果显示,多数情况下大模型评分与人类评估高度一致,可有效捕捉威胁性情绪线索。但在真实用户生成数据上的表现低于合成数据,表明对现实数据的建模仍有提升空间。文本多样性分析显示,合成数据的语义多样性略低于真实数据。总体表明,该方法在克服数据获取伦理与物流障碍的前提下,具备实现内鬼情绪检测的可行性与可扩展性。

原文摘要 · Abstract (English)

Insider threats wield an outsized influence on organizations, disproportionate to their small numbers. This is due to the internal access insiders have to systems, information, and infrastructure. %One example of this influence is where anonymous respondents submit web-based job search site reviews, an insider threat risk to organizations. Signals for such risks may be found in anonymous submissions to public web-based job search site reviews. This research studies the potential for large language models (LLMs) to analyze and detect insider threat sentiment within job site reviews. Addressing ethical data collection concerns, this research utilizes synthetic data generation using LLMs alongside existing job review datasets. A comparative analysis of sentiment scores generated by LLMs is benchmarked against expert human scoring. Findings reveal that LLMs demonstrate alignment with human evaluations in most cases, thus effectively identifying nuanced indicators of threat sentiment. The performance is lower on human-generated data than synthetic data, suggesting areas for improvement in evaluating real-world data. Text diversity analysis found differences between human-generated and LLM-generated datasets, with synthetic data exhibiting somewhat lower diversity. Overall, the results demonstrate the applicability of LLMs to insider threat detection, and a scalable solution for insider sentiment testing by overcoming ethical and logistical barriers tied to data acquisition.

内鬼检测大模型应用合成数据情感分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。