arXiv:2603.22499cs.CRcs.LG2026-03被引 4

构建可验证的合成内鬼检测数据集,解决生成文本自相矛盾问题。

OrgForge-IT: A Verifiable Synthetic Benchmark for LLM-Based Insider Threat Detection

  • 用确定性模拟引擎保证跨文档一致性,语言模型只生成表面文本。
  • 含2904条高噪声日志,覆盖四类威胁场景,验证多天关联检测能力。
  • 适合研究大模型内鬼检测、系统评估与提示工程的学者和工程师。

合成内鬼威胁数据集存在一致性缺陷:无外部事实约束生成的语料无法排除跨文档矛盾。当前领域基准CERT数据集静态且缺乏跨平台关联场景,已不适应大模型时代。本文提出OrgForge-IT,一个可验证的合成基准,通过确定性模拟引擎维护真实状态,仅由语言模型生成表面文本,确保跨文档一致性为架构保障。数据集覆盖51个模拟日,包含2,904条遥测记录,噪声率达96.4%,设计了四种检测场景,旨在击败单一表面或单日分析策略,涵盖三类威胁和八种可注入行为。十模型排行榜揭示多个发现:(1) 调查与判决准确率分离——八模型调查看似一致(F1=0.80),判决准确率却在1.0与0.80间分野;(2) 基线误报率是判决准确率的必要伴生指标,相同判决准确率下,调查看错率相差两个数量级;(3) 在语音诈骗场景中,层级A模型能排除被攻陷账户持有者,层级B模型虽检测到攻击但错误认定受害者;(4) 严格多信号阈值结构会天然排除单一表面疏忽型内鬼,凸显并行、威胁类别特异的调查管道必要性;(5) 代理式软件工程训练对多日时间关联具有增强效应,但仅在前沿参数规模下生效。最后,提示敏感性分析显示非结构化提示引发词汇幻觉,推动采用双轨评分框架区分提示遵循与推理能力。OrgForge-IT开源,采用MIT许可。

原文摘要 · Abstract (English)

Synthetic insider threat benchmarks face a consistency problem: corpora generated without an external factual constraint cannot rule out cross-artifact contradictions. The CERT dataset -- the field's canonical benchmark -- is also static, lacks cross-surface correlation scenarios, and predates the LLM era. We present OrgForge-IT, a verifiable synthetic benchmark in which a deterministic simulation engine maintains ground truth and language models generate only surface prose, making cross-artifact consistency an architectural guarantee. The corpus spans 51 simulated days, 2,904 telemetry records at a 96.4% noise rate, and four detection scenarios designed to defeat single-surface and single-day triage strategies across three threat classes and eight injectable behaviors. A ten-model leaderboard reveals several findings: (1) triage and verdict accuracy dissociate - eight models achieve identical triage F1=0.80 yet split between verdict F1=1.0 and 0.80; (2) baseline false-positive rate is a necessary companion to verdict F1, with models at identical verdict accuracy differing by two orders of magnitude on triage noise; (3) victim attribution in the vishing scenario separates tiers - Tier A models exonerate the compromised account holder while Tier B models detect the attack but misclassify the victim; (4) rigid multi-signal thresholds structurally exclude single-surface negligent insiders, demonstrating the necessity of parallel, threat-class-specific triage pipelines; and (5) agentic software-engineering training acts as a force multiplier for multi-day temporal correlation, but only when paired with frontier-level parameter scale. Finally, prompt sensitivity analysis reveals that unstructured prompts induce vocabulary hallucination, motivating a two-track scoring framework separating prompt adherence from reasoning capability. OrgForge-IT is open source under the MIT license.

内鬼检测合成数据大模型评估可验证基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。