对比四款AI安全防护工具,DKnownAI表现最佳
A Comparative Evaluation of AI Agent Security Guardrails
- 用人工标注作基准,测试各工具对两类风险的检测能力
- DKnownAI召回率96.5%,误报率90.4%,综合性能最优
- 适合关注AI代理安全、需高精度防护的开发者和企业
本报告对DKnownAI Guard在AI代理安全场景下的表现进行了对比评估,与AWS Bedrock Guardrails、Azure Content Safety及Lakera Guard三款竞品进行比较。以人工标注为真实标签,评估各防护工具对两类风险的检测能力:一是针对代理自身的威胁(如指令覆盖、间接注入、工具滥用),二是旨在诱导生成有害内容的请求(如仇恨言论、色情、暴力)。评估结果显示,DKnownAI Guard召回率达96.5%,真负率(TNR)达90.4%,在所有评测防护工具中表现最优。
原文摘要 · Abstract (English)
This report presents a comparative evaluation of DKnownAI Guard in AI agent security scenarios, benchmarked against three competing products: AWS Bedrock Guardrails, Azure Content Safety, and Lakera Guard. Using human annotation as the ground truth, we assess each guardrail's ability to detect two categories of risks: threats to the agent itself (e.g., instruction override, indirect injection, tool abuse) and requests intended to elicit harmful content (e.g., hate speech, pornography, violence). Evaluation results demonstrate that DKnownAI Guard achieves the highest recall rate at 96.5\% and ranks first in true negative rate (TNR) at 90.4\%, delivering the best overall performance among all evaluated guardrails.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。