arXiv:2601.07880cs.CRcs.AI2026-01被引 3

首个评估智能体AI在身份安全可见性任务表现的基准测试

Sola-Visibility-ISPM: Benchmarking Agentic AI for Identity Security Posture Management Visibility

  • 构建真实企业环境下的智能体系统评测框架
  • 在77个问题中实现0.84专家准确率与0.77严格成功率
  • 适用于安全研究人员和AI系统开发者参考

身份安全态势管理(ISPM)是现代企业在云与SaaS环境中面临的核心挑战。理解身份资产清单与配置健康度等基础问题,需解析复杂身份数据,推动了对智能体AI系统的需求。然而,目前尚无标准方法评估此类系统在真实企业数据上的表现。本文提出Sola Visibility ISPM基准,首个基于实际生产环境(涵盖AWS、Okta、Google Workspace)的评测体系,聚焦身份清单与健康度问题。配套的Sola AI Agent能将自然语言查询转化为可执行的数据探索步骤,生成可验证的答案。在77个测试问题中,该代理整体表现良好:专家准确率为0.84,严格成功率为0.77。在AWS健康度任务上表现最佳,专家准确率达0.94;在Google Workspace与Okta任务上结果中等但具竞争力。本工作为评估智能体AI在身份安全领域的表现提供了可复现的基准,并为未来更高级的身份分析与治理评测奠定基础。

原文摘要 · Abstract (English)

Identity Security Posture Management (ISPM) is a core challenge for modern enterprises operating across cloud and SaaS environments. Answering basic ISPM visibility questions, such as understanding identity inventory and configuration hygiene, requires interpreting complex identity data, motivating growing interest in agentic AI systems. Despite this interest, there is currently no standardized way to evaluate how well such systems perform ISPM visibility tasks on real enterprise data. We introduce the Sola Visibility ISPM Benchmark, the first benchmark designed to evaluate agentic AI systems on foundational ISPM visibility tasks using a live, production-grade identity environment spanning AWS, Okta, and Google Workspace. The benchmark focuses on identity inventory and hygiene questions and is accompanied by the Sola AI Agent, a tool-using agent that translates natural-language queries into executable data exploration steps and produces verifiable, evidence-backed answers. Across 77 benchmark questions, the agent achieves strong overall performance, with an expert accuracy of 0.84 and a strict success rate of 0.77. Performance is highest on AWS hygiene tasks, where expert accuracy reaches 0.94, while results on Google Workspace and Okta hygiene tasks are more moderate, yet competitive. Overall, this work provides a practical and reproducible benchmark for evaluating agentic AI systems in identity security and establishes a foundation for future ISPM benchmarks covering more advanced identity analysis and governance tasks.

智能体AI身份安全基准测试云安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。