arXiv:2603.14987cs.CLcs.DB2026-03

构建可衡量的智能体可信度评估体系,突破孤立测试局限。

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI

  • 定义五维度可信度:可靠、鲁棒、安全、伦理对齐、操作完整。
  • 提出分布感知采样框架,揭示单一指标无法发现的权衡关系。
  • 跨模型家族验证成功,13个系统均提升,2个达完美可信度。

智能体系统通过工具增强的多步流程执行任务,其失败(如危险工具使用、越权行为、社会危害)具有部署级后果。当前评估碎片化,集中在孤立基准片段,'可信度'常被提及却缺乏可操作定义。本文指出核心局限在于:(i) 缺乏可信度的可测量规范,(ii) 缺少原则性代表性的概念,难以在社会技术场景分布上进行评估。为此,我们提出五属性可信度定义(可靠性、鲁棒性、安全性、社会-伦理对齐、操作完整性),基于现有AI风险框架;并构建全息智能体评估框架(HAAF),通过静态策略分析、沙箱仿真、伦理对齐评估与分布感知采样,结合迭代可信优化工厂,将红队诊断转化为蓝队干预。贡献包括:(1) 可操作的五属性可信度定义;(2) 分布感知场景采样框架,揭示单标量排行榜无法察觉的属性权衡;(3) 跨家族迁移实验:从单一焦点模型设计的干预措施,无需针对每模型或场景调优,成功应用于7个模型族(Llama、Mistral、Kimi、GLM、Qwen、GPT、DeepSeek)共13个系统,在100个场景套件上实现全部系统提升,其中2个达完美风险加权可信度,证明HAAF工厂为模型无关的部署就绪管道。代码开源:https://github.com/TonyQJH/haaf-pilot

原文摘要 · Abstract (English)

Agentic AI systems increasingly act through tool-augmented, multi-step workflows whose failures (unsafe tool use, unauthorised actions, social harm) carry deployment-level consequences. Evaluation practice remains fragmented across isolated benchmark slices, and "trustworthiness" is frequently invoked but rarely defined operationally. We argue the central limitation is twofold: (i) the absence of a measurable specification of what agent trustworthiness means, and (ii) the lack of a principled notion of representativeness allowing assessment over a socio-technical scenario distribution rather than disconnected benchmark instances. We address (i) by defining agentic trustworthiness as a five-property profile (Reliability, Robustness, Safety, Social-Ethical Alignment, Operational Integrity) grounded in current AI risk frameworks, and (ii) with the Holographic Agent Assessment Framework (HAAF), which measures this profile over a scenario manifold through static policy analysis, sandbox simulation, social-ethical alignment assessment, and distribution-aware sampling, connected through an iterative Trustworthy Optimization Factory that converts red-team diagnoses into blue-team interventions. Our contributions are: (1) an operational five-property definition of agentic trustworthiness; (2) a distribution-aware scenario-sampling framework that surfaces property-level trade-offs invisible to scalar leaderboards; and (3) a cross-family transfer experiment in which interventions designed from a single focal model generalise -- without per-model or per-scenario tuning -- to 13 systems from seven model families (Llama, Mistral, Kimi, GLM, Qwen, GPT, DeepSeek) on a 100-scenario suite, where all 13 systems improve and two reach a perfect risk-weighted profile, establishing HAAF's Factory as a model-agnostic deployment-readiness pipeline. Code: https://github.com/TonyQJH/haaf-pilot

智能体评估可信度多模型迁移风险控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。